Hanbin Hong

Research Scientist · Content Safety

Hanbin Hong

I work on scaling LLM/VLM post-training for real-world content moderation through large-scale automated data annotation, data engineering, and iterative model evaluation. My research background is in AI Security, including prompt security, adversarial machine learning, and certified robustness.

Binary portrait of Hanbin Hong
Currently Research Scientist @ TikTok Content Safety · San Jose, California

What I work on

Building scalable model and data systems for Content Safety.

My current work focuses on how LLM/VLM post-training can scale through better supervision, training-data construction, difficult-example discovery, and iterative evaluation. This work builds on my broader research experience in machine learning security and robustness.

A

Post-training for Content Safety

Develop data-centric approaches for scaling LLM/VLM post-training in large-scale, long-tail content moderation scenarios.

LLM/VLM post-trainingLong-tail moderationData-centric ML
B

Automated data & supervision systems

Build large-scale automated annotation, quality-control, and data-processing systems that transform real-world content and model feedback into reliable training signals.

Automated annotationScalable supervisionData flywheels
C

AI Security research

Study prompt security, adversarial attacks and defenses, security evaluation, and certified robustness for machine learning systems.

Prompt securityAdversarial MLCertified robustness

From real-world content to continuous improvement

  1. Real-world content
  2. Automated supervision
  3. Training data at scale
  4. LLM/VLM post-training
  5. Evaluation and iteration

Featured work · PromptSecurity

My experience in systematic security evaluation informs how I think about measurement, reliability, and failure analysis in real-world model systems.

Featured research · arXiv v3 · July 2026

SoK: Systematizing LLM Prompt Security Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

A unified view of jailbreak attacks, defenses, model vulnerabilities, datasets, and evaluation. The work pairs three linked taxonomies with an auditable protocol, a modular open-source platform, two public datasets, and an interactive leaderboard.

One auditable experiment tuple
Dataset
Attack
Defense
Target LLM
Output
Judger

Compare compatible configurations while preserving settings, costs, outputs, and sample-level traces.

445,752 jailbreak prompt pairs from 48 sources
1.09M benign prompts from 14 sources
175,400 evaluation records harmful + benign utility
11 · 21 · 10 models · attack settings · defense states with 3 judgers
Release figures from arXiv v3 · Dataset access: JailbreakDB ↗ PromptSecurity-Eval ↗ Direct PDF ↗

Selected findings

Security conclusions depend on what—and how—you measure.

Unified evaluation exposes failure modes that single-number rankings often miss: worst-case attacks, safety–utility tradeoffs, and sensitivity to the chosen judger.

41.8%
Worst-case exposure

Averages can hide concentrated risk.

GPT-5.2 showed 4.7% average induced harmfulness, while its worst-case ASR under GPTFUZZER reached 41.8%.

−60.6 pp
Safety–utility tradeoff

Mitigation strength is not the whole story.

Output filtering reduced defended ASR to 2.7%, but benign-task accuracy fell by 60.6 percentage points.

71–87 pp
Measurement sensitivity

The judger can change the conclusion.

For identical model outputs, changing only the HarmBench behavior input shifted measured ASR by 71–87 points.

Selected publications

AI security and robustness.

Selected work across prompt security, certifiable attacks, textual robustness, and randomized smoothing. Visit Google Scholar for the complete list.

Towards Strong Certified Defense with Universal Asymmetric Randomization

IEEE CSF · 2026

Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence

ACM CCS · 2024

Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks

IEEE S&P · 2024

UniCR: Universally Approximated Certified Robustness via Randomized Smoothing

ECCV · 2022

View all publications on Google Scholar

About and background

A research foundation for reliable model systems.

Across industry and academia, I am interested in how better data, supervision, and evaluation can make machine learning systems more reliable in practice.

Current work and research background

I am a Research Scientist at TikTok working on Content Safety, with a focus on scaling LLM/VLM post-training through large-scale automated data annotation, data engineering, and model evaluation.

My research background spans prompt security, adversarial machine learning, and certified robustness. Across these areas, I am interested in how better data, supervision, and evaluation can make machine learning systems more reliable in practice.

Contact

Let’s connect.

I am open to conversations about LLM/VLM post-training, automated data annotation, data-centric model development, Content Safety, prompt security, and adversarial machine learning.