I study how AI systems can keep learning over time without destroying what they
already know, something humans do effortlessly. I work on continual learning and
meta-learning, in diffusion models and large language models.
Continual learning. I work on consolidating knowledge into a model
as it keeps learning. External stores work well, but I believe some of what a model
learns has to live in the model itself.
Meta-learning. Today, every decision in training — what to learn
from, what to change, how much — is made for the model. I work towards models
that make these decisions themselves, starting with structured autonomy:
the model makes some of them, within a structure we design.
The two are closely related. The most natural way to consolidate is through what the
model already knows — self-distillation or RL in LLMs, or composing learned
knowledge in diffusion models. Such updates diverge less from the model, so they
forget less, and they hand more of the learning to the model itself.
I also love to teach, and have been a TA for four semesters. Grading in large
courses takes TAs a long time, so I have worked on training models to give students
feedback in the meantime.
Trust Region Continual Learning as an Implicit Meta-Learner was accepted
to NeurIPS 2026.
CobwebTM was accepted to Findings of ACL 2026.
Avoid Catastrophic Forgetting with Rank-1 Fisher was accepted to
ICLR 2026.
Gave an oral presentation on Hierarchical Semantic Retrieval with Cobweb
at
ACS 2025.
Research
Continual learningMeta-learning
Continual Learning in Diffusion Models
with Christopher MacLellan
We show that the empirical Fisher of a diffusion model is rank-1 in low-SNR regimes.
Using this, we define a rank-1 EWC penalty and show that, combined with replay, the
model forgets less. We then show that this formulation is equivalent to one-step
MAML under local approximations, and observe the effect directly: the model
re-learns previous tasks faster than other methods.
Constraining how much you update is already a form of learning to
learn.
The same intuition as the diffusion work, but at LLM scale the Fisher is too
expensive to compute, so SCoL lets the model learn it: meta-RL decides which of
its own layers to change when it absorbs a passage. Trained only on short
contexts, it beats in-context baselines that see the whole passage on LongBench
v2.
Consolidation with autonomy: the model decides where new knowledge
goes.
Cobweb builds a concept hierarchy incrementally, one example at a time. As a
retrieval index it matches dense encoders and stays robust where kNN collapses;
CobwebTM extends it to lifelong hierarchical topic modeling and matches or beats prior
incremental methods.
The external-memory side: a retrieval structure that is itself
continual.
One language model is trained in several study roles over the same examples —
comparing problems, identifying subgoals, explaining. Roles that never produce an
answer still improve answer accuracy.
Roles hand more of the training to the model itself; next, the model
designs its own roles.
Preprint
Structured autonomy
Test-Time Compositional Generation
with Christopher MacLellan
From a single out-of-distribution query, recover concept prototypes by score-based
mode finding across the noise levels of a pretrained diffusion model, then compose
them into a product-of-experts teacher.
Once the algorithm is set, the only open question is what to
incorporate.
Given a specific misconception, can a model produce the wrong answer a real student
holding it would produce, in a single turn? A counterfactual generation problem.
Human-like errors can benchmark and train models; here, for grading
theoretical computer science, where partial credit is hard.
Selected publications
* denotes equal contribution.
NeurIPS2026
Trust Region Continual Learning as an Implicit Meta-Learner
Zekun Wang*, Anant Gupta*, Christopher J. MacLellan
Conference on Neural Information Processing Systems · Main Track ·
Poster · Atlanta, Dec 2026
Four semesters of TAing Automata and Complexity shaped how I think about learning in
both humans and machines — and led directly to my work on LLMs for education.
Jan 2024 – Dec 2025
Teaching Assistant — CS 4510: Automata and Complexity
Georgia Institute of Technology · Head TA, Spring 2025
Delivered two lectures to 300+ students, led 12 review sessions, and wrote
homework and exam problems.
Jan 2023 – Aug 2023
College of Computing Tutor
Georgia Institute of Technology
Tutored 150+ hours across AI, algorithms, data structures, and discrete
mathematics.
Service
Reviewer, International Conference on Learning Representations (ICLR 2027)
Reviewer, NeurIPS 2026 Workshop on Towards Test-Time Continual Learning
Agents (TTCL)
Education
M.S. Computer Science, Georgia Tech — expected May 2027 GPA 4.0/4.0
I'm always happy to talk about continual learning, meta-learning, or LLMs for
education — and I'm glad to hear from students thinking about research. The
fastest way to reach me is email.