Post-training for Content Safety
Develop data-centric approaches for scaling LLM/VLM post-training in large-scale, long-tail content moderation scenarios.
Research Scientist · Content Safety
I work on scaling LLM/VLM post-training for real-world content moderation through large-scale automated data annotation, data engineering, and iterative model evaluation. My research background is in AI Security, including prompt security, adversarial machine learning, and certified robustness.
What I work on
My current work focuses on how LLM/VLM post-training can scale through better supervision, training-data construction, difficult-example discovery, and iterative evaluation. This work builds on my broader research experience in machine learning security and robustness.
Develop data-centric approaches for scaling LLM/VLM post-training in large-scale, long-tail content moderation scenarios.
Build large-scale automated annotation, quality-control, and data-processing systems that transform real-world content and model feedback into reliable training signals.
Study prompt security, adversarial attacks and defenses, security evaluation, and certified robustness for machine learning systems.
From real-world content to continuous improvement
Featured work · PromptSecurity
My experience in systematic security evaluation informs how I think about measurement, reliability, and failure analysis in real-world model systems.
A unified view of jailbreak attacks, defenses, model vulnerabilities, datasets, and evaluation. The work pairs three linked taxonomies with an auditable protocol, a modular open-source platform, two public datasets, and an interactive leaderboard.
Compare compatible configurations while preserving settings, costs, outputs, and sample-level traces.
Taxonomies, protocol, experiments, and findings.
Open on arXivBrowse compatible main-protocol measurements.
Explore resultsA consolidated resource for jailbreak and benign prompts.
View on Hugging FacePromptSecurity-Eval harmfulness and utility results.
View on Hugging FaceCompose models, attacks, defenses, data, and judgers.
View on GitHubSelected findings
Unified evaluation exposes failure modes that single-number rankings often miss: worst-case attacks, safety–utility tradeoffs, and sensitivity to the chosen judger.
GPT-5.2 showed 4.7% average induced harmfulness, while its worst-case ASR under GPTFUZZER reached 41.8%.
Output filtering reduced defended ASR to 2.7%, but benign-task accuracy fell by 60.6 percentage points.
For identical model outputs, changing only the HarmBench behavior input shifted measured ASR by 71–87 points.
Selected publications
Selected work across prompt security, certifiable attacks, textual robustness, and randomized smoothing. Visit Google Scholar for the complete list.
arXiv:2510.15476 · revised July 2026
IEEE CSF · 2026
ACM CCS · 2024
IEEE S&P · 2024
ECCV · 2022
About and background
Across industry and academia, I am interested in how better data, supervision, and evaluation can make machine learning systems more reliable in practice.
I am a Research Scientist at TikTok working on Content Safety, with a focus on scaling LLM/VLM post-training through large-scale automated data annotation, data engineering, and model evaluation.
My research background spans prompt security, adversarial machine learning, and certified robustness. Across these areas, I am interested in how better data, supervision, and evaluation can make machine learning systems more reliable in practice.
Contact
I am open to conversations about LLM/VLM post-training, automated data annotation, data-centric model development, Content Safety, prompt security, and adversarial machine learning.