Ethan (Yi) Zeng

Safety Research Lead, Foundation Model · TikTok

I lead foundation-model safety research at TikTok, working to make safety a capability models learn and retain. Together with research, engineering, and policy teams, I build the data, training, and evaluation systems that put this research into practice.

Ph.D. in Computer Engineering, Virginia Tech, advised by Ruoxi Jia.

Yi Zeng with a dog

Research perspective

Safety as a learned capability

Our earlier work showed that fine-tuning can undo alignment. I study safety across the full model lifecycle: how it is learned from data during pretraining and mid-training, shaped by post-training, and retained through adaptation. Evaluation and deployment feedback guide the next round of data and training.

Experience at TikTok 2025–present
Safety Research Lead, Foundation ModelSafety training & evaluation, including agents and red-teaming
Senior Staff Research ScientistPlatform-wide data flywheel for model training & evaluation
Staff Research ScientistAlignment data & deliberative alignment across regional policies

Current scope

Research I lead at TikTok

Data & Pretraining

Platform data is organized and curated for model training and evaluation; capability gaps guide the next data iteration.
Concept illustration

2026 / Platform-wide data infrastructure

TikTok data flywheel

Architected TikTok's data flywheel and lead its development with research and engineering partners.

Details

The flywheel connects platform-wide data with foundation-model training and evaluation. Shared infrastructure supports data discovery, curation, and coverage analysis for both general capabilities and safety; evaluation gaps guide the next round of data development.

A platform-wide data flywheel supporting general model capabilities and safety. Concept illustration, without specific data sources, implementation details, or measured results.

General knowledge and safety reasoning are considered together in a model-learning research question.
Concept illustration

2026 / Unpublished research

Safety capability pretraining

Proposed the safety-pretraining direction and led its development with research collaborators.

Details

I study how safety understanding can develop alongside general knowledge. This research treats pretraining as a place to learn safety capabilities, and evaluation as a way to distinguish those capabilities from refusal alone.

The research question of learning safety alongside knowledge. This is not a training recipe or an experimental result.

Post-training & Adaptation

An alpha-divergence dial illustrates the choice between mode-seeking and mode-covering in post-training.
Concept illustration

2026 / NeurIPS, accepted

A dial for post-training

Co-developed an α-divergence dial for mode-seeking and mode-covering in post-training.

Details

I co-developed an α-divergence dial for post-training: choose how strongly learning concentrates on a few high-reward responses or covers a broader range. Tests across reasoning, agentic tasks, and safety show that the useful setting depends on the task and reference model.

A conceptual view of post-training geometry, not an original paper figure or a universally optimal setting. Conference listing

A model encounters new tasks, raising the question of which learned capabilities remain after adaptation.
Concept illustration

2026 / Unpublished research

Capability retention during adaptation

Proposed and co-developed a training-time method for capability retention.

Details

Earlier work on fine-tuning safety raised a practical question: how can a model learn new tasks without losing useful behavior? I proposed the method, co-developed it with a research intern, and directed the experimental design.

The question of capability retention during adaptation. No method implementation or measured outcome is depicted.

Evaluation & Red-teaming

Risk understanding, safe responses, and safe actions are connected, with robustness across contexts spanning all three.
Concept illustration

2026 / Evaluation research

Pan-safety evaluation & scaling

Designed evaluations spanning risk understanding, safe responses, and agent behavior.

Details

Refusal scores capture only part of safety. Pan-Safety connects risk understanding, generation, agent behavior, and robustness in a shared evaluation framework. We also study whether signals from mid-training checkpoints can predict how training approaches rank on safety behavior after post-training, so training decisions can be informed earlier.

Connected research dimensions, not model scores, institutional comparisons, or an exhaustive safety taxonomy.

Attack proposals are tested on a model, independently checked, and refined through feedback.
Concept illustration

2026 / Adversarial evaluation

Red-teaming models & evaluation

I lead the training of dedicated in-house red-teaming models.

Details

Our post-training work focuses on increasing the models' propensity to generate attacks and improving attack quality across diverse target models. Evaluation checks whether those attacks reveal genuine target-model failures.

A functional red-teaming loop. The illustration does not encode attack results.

Preparedness & Deployment

Misuse, influence, and agent scenarios call for distinct response, behavior, and action evaluations.
Concept illustration

2026 / Evaluation infrastructure

Preparedness & agentic evaluation

Built model-evaluation and risk-review workflows with research, policy, and legal colleagues.

Details

My work spans misuse, persuasion-related behavior, political bias, and agentic evaluation. Each asks a different measurement question; a model response, an action in an environment, and an effect on a person are not interchangeable evidence.

Distinct measurement questions inform risk review. The illustration names no data sources and reports no results.

Safety requirements guide training data, model alignment, and release evidence, with safety and capability checks.
Concept illustration

2025 / Production alignment

Global model safety alignment

Technical lead for global LLM/VLM safety alignment across training, policy, and evaluation.

Details

I connected regional safety requirements to alignment data and model evaluation, including work on bias mitigation and deliberative alignment. That production experience shaped my current focus on how safety is learned throughout training.

The alignment workflow, not a claim of universal safety or legal compliance.

Published research

Google Scholar

Selected publications

AutoRedTeamer combines red-teaming evaluation with a separate attack-strategy integration loop.

AutoRedTeamer

NeurIPS

A red-teaming agent generates and runs tests while a second agent integrates new attacks from research. Memory-guided selection lets the system reuse what it learns as threats evolve.

A risk-taxonomy overview comparing the categories covered by AIR-Bench and earlier benchmarks.

AIR-Bench 2024

ICLR / Spotlight

Turns risks specified in government regulations and company policies into a shared taxonomy and safety benchmark, making gaps in model coverage easier to compare across jurisdictions.

Frequency of safety categories in ten prior datasets, ordered from most to least represented.

SORRY-Bench

ICLR

Balances unsafe topics and varies how requests are phrased to make refusal evaluation more systematic. Human annotations also support smaller, less costly automated judges.

A persuasion taxonomy guides query paraphrasing into human-readable adversarial prompts, which are tested on language models.

How Johnny Can Persuade LLMs to Jailbreak Them

ACL / Best Social Impact Paper

Uses social-science persuasion techniques to construct readable jailbreaks. These attacks expose weaknesses in safeguards that can miss harmful intent when it is wrapped in convincing language.

Conflicting group-wise hypergradients motivate a two-stage meta-learning method using Nash bargaining before fairness optimization.

Fairness-Aware Meta-Learning via Nash Bargaining

NeurIPS

Treats conflicts between demographic groups' learning signals as a bargaining problem. A two-stage method first reconciles those signals, then optimizes the chosen fairness objective.

Earlier publications 2021–2023