A taxonomy-guided collection of research on specifying, supervising, implementing, and assuring Human–AI Alignment.
This survey provides a unified overview of Human–AI Alignment. We propose a lifecycle taxonomy that organizes the field along four dimensions—alignment specification, supervision, mechanisms, and assurance—and use it to structure the literature across 23 non-exclusive research branches.
Reference components for 22 cited papers in Task & Assistance, Personalized Alignment, and Uncertainty & Drift are documented in Alignment specification, with explicit implementation limits and an offline demo.
Human Feedback and AI Feedback code, covering the supervision components of 17 cited papers, is documented in Alignment supervision.
The repository also provides a Python framework that mirrors the four lifecycle dimensions while keeping common workflows simple. The core has no runtime dependencies; training libraries are installed only for the methods that need them.
python -m pip install -e .
hai-align catalog validate
python examples/minimal_pipeline.py
To train with DPO:
python -m pip install -e ".[dpo]"
hai-align train dpo --model Qwen/Qwen3-0.6B --dataset trl-lib/ultrafeedback_binarized --output outputs/qwen-dpo
from human_alignment import DPO
run = DPO(
model="Qwen/Qwen3-0.6B",
dataset="trl-lib/ultrafeedback_binarized",
output_dir="outputs/qwen-dpo",
).train()
print(run.generate("What is Human--AI alignment?"))
See Codebase architecture for the public API and extension points. The machine-readable taxonomy lives in src/human_alignment/catalog/data/taxonomy.json.
Preference distillation is available through one consistent API for VPD, PPD, DCKD, TVKD, ADPA, and CTPD:
python -m pip install -e ".[distillation]"
from human_alignment import VPD, PreferenceDistillationExample
data = [
PreferenceDistillationExample(
prompt="Explain alignment briefly.",
responses=("Alignment connects behavior to human targets.", "It is model scaling."),
teacher_scores=(1.0, 0.0),
)
]
run = VPD(model="student-model", dataset=data, output_dir="outputs/student-vpd").train()
See Preference distillation for objective-specific dataset schemas, teacher-model use, and migration details.
The benchmarks the survey uses as evidence (RewardBench 2, JudgeBench, AlpacaEval 2) run under their official scoring rules, and three inference-time methods (CAA, best-of-N with a reward model, ARGS) steer a frozen model:
hai-align bench rewardbench2 --reward-model Skywork/Skywork-Reward-V2-Qwen3-0.6B --output-dir eval/rb2
See Benchmarks and inference-time methods.
Safety-alignment training is integrated under the same package, with 25 stages covering SafeRLHF, SafeDPO, BSO, SACPO, CAN, MODPO, CPO, BFPO, MidPO, reward and cost modeling, and multi-objective RLHF:
python -m pip install -e ".[safety]"
hai-align safety list
hai-align safety run recipes/safety_alignment/methods/safedpo.yaml
from human_alignment import SafetyAlignment
run = SafetyAlignment(
config="recipes/safety_alignment/methods/safedpo.yaml",
output_dir="outputs/safedpo",
).train()
See Safety alignment for recipe dependencies, paper-faithful implementation choices, evaluation, and H100 workflows.
Additional preference methods are available as IPO, BPO, TDPO, TISDPO,
TIDPO, TBPOQ, and TBPOA, with recipes under recipes/preference_optimization/.
See Preference optimization for installation,
token-weight data formats, objective sources, and checkpoint handling.
The library supports both ends of an alignment experiment:
from human_alignment import DPO, load_checkpoint
# Fine-tune a pretrained model or an existing checkpoint.
trained = DPO(
model="models/student_sft",
dataset="data/preferences",
output_dir="outputs/student_dpo",
).train()
# Load the resulting checkpoint later for inference or assurance.
checkpoint = load_checkpoint(
"outputs/student_dpo",
config={"device_map": "auto", "torch_dtype": "bfloat16"},
)
print(checkpoint.generate("Explain alignment."))
Preference-distillation supervision can be prepared without the temporary research scripts:
hai-align prepare distillation dckd \
--dataset HuggingFaceH4/ultrafeedback_binarized \
--teacher models/teacher_dpo \
--tokenizer models/student_sft \
--output data/ultrafeedback-dckd
| Example | Purpose |
|---|---|
minimal_pipeline.py | Validate the catalog and run a minimal end-to-end pipeline |
dpo_quickstart.py | Launch DPO from Python |
dpo.toml | Configure a DPO run declaratively |
checkpoint_workflow.py | Train, reload, and use a checkpoint |
preference_distillation.py | Prepare and run preference distillation |
Human–AI alignment is organized as a lifecycle spanning alignment specification, supervision, mechanisms, and assurance across chat, code, mathematical reasoning, multimodal, agentic, robotics, and healthcare settings.
View the high-resolution survey overview (PDF)
The lifecycle taxonomy organizes Human–AI alignment into four dimensions and eight groups; the paper collection below further resolves them into 23 non-exclusive research branches.
View the high-resolution taxonomy (PDF)
| Dimension | Guiding question | Groups |
|---|---|---|
| Alignment Specification | What should an AI system align to, and whose objectives and values should count? | Alignment Objectives · Values & Stakeholders |
| Alignment Supervision | Where do alignment signals come from, and how are they expressed and scaled? | Feedback Source · Feedback & Oversight |
| Alignment Mechanisms | How are alignment signals translated into model behavior during training and inference? | Training-Time Alignment · Inference-Time Alignment |
| Alignment Assurance | How do we evaluate, stress-test, preserve, interpret, and monitor alignment? | Evaluation & Robustness · System Assurance |
Training-time alignment covers reward and verifier modeling, supervised alignment, preference optimization, reinforcement learning, and alignment distillation.
View the high-resolution training-time diagram (PDF)
Inference-time alignment covers steering, search, iterative refinement, and interactive or agentic control while keeping the model policy fixed.
View the high-resolution inference-time diagram (PDF)
Feedback granularity and optimization granularity are tracked separately because the unit receiving feedback can differ from the unit optimized by the learning objective.
Papers are sorted by year within each branch. Each entry links to the publication page, DOI, arXiv record, or a clearly labeled Scholar search when the BibTeX record has no direct link.
What should an AI system align to, and whose objectives and values should count?
| Group | Branch | Papers |
|---|---|---|
| Alignment Objectives | Task & Assistance Alignment | 35 |
| Alignment Objectives | Safety Alignment | 27 |
| Values & Stakeholders | Personalized Alignment | 24 |
| Values & Stakeholders | Pluralistic & Societal Alignment | 35 |
| Values & Stakeholders | Context, Uncertainty & Drift | 23 |
purpura2026mosaic · DOIkwon2026reasonif · DOIdong2026ifevalpp · DOIrobinette2026verify · DOIxiao2026sftmix · Paperli2025fbbench · Paperzhang2025iopo · Paperzhang2025longreward · Paperharada2025manyifeval · Paperbalepur2025goodplan · Paperxu2025magpie · Paperan2025ultraif · Paperliu2024llavarlhf · Paperwang2024helpsteer2 · DOIyu2024lions · Papersun2024parrot · Paperniu2024ragtruth · Paperchung2024flan · Papercui2024ultrafeedback · Paperrafailov2023dpo · Papergao2023citations · Paperwu2023finegrained · Paperzhou2023lima · Paperwang2023selfinstruct · Paperdong2023steerlm · Paperlongpre2023flancollection · Papermishra2022natural · Paperwei2022flan · Paperglaese2022sparrow · Papersanh2022t0 · Paperwang2022supernatural · Paperbai2022hh · Paperouyang2022instructgpt · Paperaskell2021assistant · Papernakano2021webgpt · Paperzhang2026llmva · Paperjiang2026els · Paperwang2026hiar · Paperin2026alttrain · Paperbrito2026safetyisnotuniversal · DOIzhang2026alphaalign · Paperyang2026lasa · Paperniu2026nspo · Paperzhang2026safetyoneshot · Paperkim2026safedpo · Paperzhang2025controllablesafetyalignment · Paperkaraman2025porover · Paperli2025explicitsafety · Papergao2025abd · Paperzhang2025constrainedalignment · Paperhuang2025antidote · Paperli2025larf · Paperwang2025lifelong · Paperguan2024deliberative · Paperzou2024circuitbreakers · Papermu2024rulebased · Scholardai2024saferlhf · Paperhuang2024vaccine · Paperji2023beavertails · Paperbai2022constitutional · Paperbai2022hh · Paperaskell2021assistant · Papergarbacea2026personalizedbenchmark · Papersun2026fpps · Paperli2026alignx · Papertan2026profilepeft · Paperma2026cbpo · Paperwu2025interaction · Paperkim2025drift · Paperthonet2025fast · Paperbu2025cope · Papercho2025ticl · Paperbalepur2025personas · Paperguan2025personalizedsurvey · Paperaroca2025prose · Paperchen2025pal · Paperzhu2025personality · Paperpadurean2025useralign · Paperchen2025pad · Paperzhang2025personajudge · Paperzhao2025rlpa · Papertan2024oppu · Paperli2024personalized · Paperpoddar2024personalizing · Paperkirk2024prism · Paperjang2023personalizedsoups · Paperpooledayan2026overtonbench · Papertan2026subgroupvalues · Paperzhang2026cultivating · Papersun2026cuma · Paperzhou2026disalign · Papernie2026perspectra · Paperkim2026valueflow · Paperzhang2026crosscultural · Paperbhattacharyya2026alpha · DOIzhu2026dvmap · Paperpan2026fgdalign · DOIzhang2026culturemanager · Paperali2026pluralistic · DOIengelmann2025pluralisticcompatibility · Papermentxaka2025democracy · Papershirali2025heterogeneous · DOIzhang2025pluralisticcot · Paperki2025culturaldebate · Paperhalpern2025pairwise · Paperchen2025pal · Paperyao2025gdpo · Paperxu2025culturespa · Paperadams2025steerable · DOIgreenblatt2024control · Paperhuang2024collective · Paperli2024culturellm · Papersiththaranjan2024distributional · Scholarfeng2024modular · Papersorensen2024pluralistic · Scholarconitzer2024socialchoice · Papermu2024rulebased · Scholarkirk2024prism · Paperbai2022constitutional · Paperbakker2022finetuning · Paperawad2018moral · DOIfang2026aam · DOIlin2026activedpo · Paperchen2026cfa · DOIxu2026fadpo · Paperkeswani2026moralchange · Paperwu2026rmrouting · Papershen2025activerm · Paperfeng2025pilaf · Paperxu2025robustdpo · Paperli2025uipo · Paperxu2025uncertaintyjudge · Paperzhang2025copr · Paperwang2025lifelong · Papersun2025ugda · Papermuldrew2024active · Scholarsiththaranjan2024distributional · Scholarboerstler2024stability · Paperkong2024perpcorrect · Paperhadfieldmenell2017ird · Scholarhadfieldmenell2017offswitch · DOIhadfieldmenell2016cirl · Scholarabbeel2004apprenticeship · Scholarng2000irl · ScholarWhere do alignment signals come from, and how are they expressed and scaled?
| Group | Branch | Papers |
|---|---|---|
| Feedback Source | Human Feedback | 20 |
| Feedback Source | AI Feedback | 23 |
| Feedback Source | Programmatic & Verifiable Feedback | 22 |
| Feedback & Oversight | Demonstrations & Preferences | 24 |
| Feedback & Oversight | Critique, Process & Trajectory Feedback | 19 |
| Feedback & Oversight | Reliable & Scalable Oversight | 25 |
shi2026wildfeedback · Papershaikh2025ditto · Paperjung2025bco · Papermiranda2025hybrid · Paperzhang2025mmrlhf · Paperliu2024llavarlhf · Paperxu2024finegrained · Paperhou2024chatglmrlhf · Paperwang2024helpsteer2 · DOIyu2024rlhfv · Paperlloret2024alt · Paperwang2023helpsteer · arXivkopf2023openassistant · arXivwu2023finegrained · Paperglaese2022sparrow · Paperbai2022hh · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperziegler2019finetuning · Paperchristiano2017preferences · Paperyin2026sao · Papercook2026checklist · Paperxu2025magpie · Paperyu2025daif · Paperye2025conj · Paperxiong2025llavacritic · DOImcaleese2024critics · Paperliu2024dlma · Paperlee2024rlaif · Paperyuan2024selfrewarding · Scholarcui2024ultrafeedback · Paperli2024vlfeedback · Paperding2023ultrachat · arXivliu2023geval · arXivzhu2023judgelm · arXivzheng2023mtbench · Papermukherjee2023orca · arXivshinn2023reflexion · Scholarsun2023selfalignment · Scholarwang2023selfinstruct · Papermadaan2023selfrefine · Paperxu2023wizardlm · arXivbai2022constitutional · Papersu2026rewardbridge · Paperyuan2026k2v · Paperzhang2025openprm · Paperdeepseek2025r1 · Paperzhang2025genrm · Papersetlur2025rewarding · Papergehring2025rlef · Paperdong2025autoif · Paperlightman2024verify · Papersingh2024restem · Paperxin2024deepseekprover15 · DOIshao2024deepseekmath · Paperwang2024mathshepherd · Papermu2024rulebased · Scholarhosseini2024vstar · Paperchen2023codet · Paperyang2023leandojo · Paperliu2023rltf · Paperli2022coderl · Paperzelikman2022star · Paperchen2021humaneval · arXivhendrycks2021apps · arXivshi2026wildfeedback · Papermiranda2025hybrid · Papershaikh2025ditto · Papertan2025pugc · Paperjung2025bco · Paperxiao2025sweetspot · Paperli2025joint · Paperxu2024finegrained · Paperchung2024flan · Papermuldrew2024active · Scholarli2024demoreward · Paperwang2024helpsteer2 · DOIyu2024rlhfv · Papercui2024ultrafeedback · Paperwu2023finegrained · Paperzhou2023lima · Paperwang2023selfinstruct · Paperlongpre2023flancollection · Papermishra2022natural · Paperwei2022flan · Paperwang2022supernatural · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperchristiano2017preferences · Paperxi2026critiquerl · Paperfan2026agentrrm · Paperzhang2026biprm · Paperyin2025dgprm · Paperzhang2025openprm · Papersetlur2025rewarding · Paperyu2025criticrm · Paperxie2025ctrl · Paperyu2025rco · Papergou2024critic · Scholarlightman2024verify · Papermcaleese2024critics · Paperyu2024rlhfv · Papercui2024ultrafeedback · Paperwu2023finegrained · Papershinn2023reflexion · Scholarmadaan2023selfrefine · Paperbai2022constitutional · Paperuesato2022process · Paperranaldi2026agentic · Paperlang2026selective · DOIyin2026partitioned · Paperqiu2026peerprediction · Paperrahman2025aidebate · Scholarnayebi2025barriers · Paperlang2025debate · Papergoel2025oversight · Paperhammond2025neuralproofs · Paperlambert2025rewardbench · Paperengels2025scaling · Papermuldrew2024active · Scholargreenblatt2024control · Paperkhan2024debate · Papersun2024easyhard · Paperkenton2024scalableoversight · Paperkirchner2024proververifier · Paperwang2024rmquality · DOIcheng2024rime · Scholarzhang2024selm · Paperburns2024weakstrong · Paperwu2023finegrained · Papergao2023overoptimization · Paperirving2018debate · Paperchristiano2018amplification · PaperHow are alignment signals translated into model behavior during training and inference?
| Group | Branch | Papers |
|---|---|---|
| Training-Time Alignment | Reward & Verifier Modeling | 27 |
| Training-Time Alignment | Supervised Alignment | 22 |
| Training-Time Alignment | Preference Optimization | 50 |
| Training-Time Alignment | Reinforcement Learning | 32 |
| Training-Time Alignment | Alignment Distillation | 9 |
| Inference-Time Alignment | Steering, Search & Refinement | 31 |
| Inference-Time Alignment | Interactive & Agentic Control | 17 |
liu2026openrubrics · Paperwang2026rationale · Paperzhou2026prismrm · Paperchen2026rmr1 · Paperzhang2026biprm · Papermiao2026adajudge · Paperpeng2025agenticrm · Paperyin2025dgprm · Paperzhang2025genrm · Paperwang2025gram · Paperxiong2025llavacritic · DOIguo2025rrm · DOIchen2025pal · Paperwang2024helpsteer2 · DOIlightman2024verify · Paperwang2024mathshepherd · Paperkirchner2024proververifier · Papercoste2024ensembles · Paperdai2024saferlhf · Paperwu2023finegrained · Paperglaese2022sparrow · Paperuesato2022process · Paperbai2022hh · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperziegler2019finetuning · Paperchristiano2017preferences · Paperlee2026oasis · Paperxiao2026sftmix · Paperxu2025magpie · Paperyang2025main · Paperharada-etal-2025-massive · Paperhua-etal-2025-intuitive · Paperzhang2025grape · Scholarxia2024less · Paperyang2024metaaligner · DOIchung2024flan · Paperliu2024selectit · Paperli2024backtranslation · Paperliu2024deita · Paperzhou2023lima · Paperwang2023selfinstruct · Paperdong2023steerlm · Paperlongpre2023flancollection · Papermishra2022natural · Paperwei2022flan · Papersanh2022t0 · Paperwang2022supernatural · Paperouyang2022instructgpt · Paperle2026causaldpo · Paperzhu2026rappo · Paperyuan2026hypo · Paperliu2026statisticalalignment · DOIoi2026autoregressive · Papernguyen2026bso · DOIkim2026safedpo · Paperluo2026sharpness · Paperyang2026token · Papernguyen2026tokenratio · Paperzhang2025constrainedalignment · Paperwu2025alphadpo · Paperdoosterlinck2025apo · Paperxu2025doublyrobust · Paperlee2025kl · Paperwang2025mpo · Paperzhu2025mmedpo · Paperyao2025gdpo · Paperkim2025bregmanpo · Paperguo2025pro · Papersun2025rpo · Paperzhang2025risk · Paperkong2025sdpo · Paperwu2025sppo · Paperzhu2025sgdpo · Paperzhou2025treg · Paperliu2025tisdpo · Paperwu2025drdpo · Paperzhu2025wspo · Papermuldrew2024active · Scholarxu2024behaviorbpo · Paperxu2024cpo · Paperrosset2024dno · Paperethayarajh2024kto · Paperyang2024metaaligner · DOIshani2024multiturn · Scholarmunos2024nash · Paperhong2024orpo · Paperchowdhury2024rdpo · Papercheng2024rime · Scholarding2024sail · Paperzhang2024selm · Papermeng2024simpo · Paperzeng2024tdpo · Paperwu2024beta · Paperazar2023psipo · Paperzhao2023hadpo · Paperrafailov2023dpo · Paperyuan2023rrhf · Paperzhao2023slichf · Paperzhang2026alphaalign · Paperjiang2026rlvrr · Paperniu2026nspo · Paperzhang2026corewarding · Paperchu2026gpg · Paperzeng2026rlve · arXivdeng2026srl · Paperviswanathan2025rlcf · Paperliu2025dapo_advantage · Paperdeepseek2025r1 · Paperpeng2025repo · Paperzhang2025mmrlhf · Paperwalder2025passk · Paperliu2025saturn · Paperguo2025spo · Paperfang2025serl · Paperleroux2025tapered · Paperliu2024llavarlhf · Paperahmadian2024reinforce · Paperhou2024chatglmrlhf · Papershao2024deepseekmath · Papershani2024multiturn · Scholarli2024remax · Paperlee2024rlaif · Papermu2024rulebased · Scholardai2024saferlhf · Paperding2024sail · Paperkaufmann2023survey · Paperuesato2022process · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperchristiano2017preferences · Papernguyen2026ctpd · Papergao2025adpa · Paperzhang2025aligndistil · Papergu2025pad · Paperkwon2025tvkd · Paperhong2024cyclealign · Paperliu2024dlma · Paperli2024dpkd · Paperzhang2024plad · Paperli2026trainingfree · arXivnguyen2026beyondlinear · arXivwu2026cybercorrect · arXivluo2026dontlosefocus · arXivzhao2026odesteer · arXivyou2026spherical · arXivnguyen2026pidsteering · Paperweng2026finesteer · Paperzhou2026requer · Papersnell2025scaling · Papershi2025ssr · arXivli2025unleashing · arXivding2025w2s · arXivmasters2025arcane · Paperzhang2025controllablesafetyalignment · Paperstolfo2025steering · Paperxiong2025llavacritic · DOIfei2025nudging · Paperhung2025darwin · Paperma2025s2r · Papermudgal2023controlled · Paperkhanov2024args · Papergou2024critic · Scholarwang2024mathshepherd · Paperrimsky2024caa · DOIturner2023activationaddition · Paperli2023iti · DOIshinn2023reflexion · Scholarzou2023repreng · Papermadaan2023selfrefine · Papernakano2021webgpt · Paperzhou2026steering · arXivzhu2026oversight · Papersuri2026clarification · Paperayankoya2025scalable · DOIshirali2025burden · arXivhudson2025corrigibility · Paperhe2025aria · Paperxue2025agsa · Papersouth2025delegation · Papergarber2025offswitch · Papergreenblatt2024control · Paperandukuri2024stargate · Scholarmozannar2020defer · Scholarhadfieldmenell2017offswitch · DOIhadfieldmenell2016cirl · Scholarorseau2016interruptible · Scholarsoares2015corrigibility · PaperHow do we evaluate, stress-test, preserve, interpret, and monitor alignment?
| Group | Branch | Papers |
|---|---|---|
| Evaluation & Robustness | Behavioral & Evaluator Evaluation | 26 |
| Evaluation & Robustness | Adversarial & Distribution Robustness | 30 |
| Evaluation & Robustness | Alignment Preservation | 17 |
| System Assurance | Mechanistic & Theoretical Evidence | 31 |
| System Assurance | Monitoring & Auditing | 29 |
wang2026planrewardbench · DOIwen2026ifrewardbench · Paperlaban2026lost · Paperxiong2026multicrit · Papermalik2026rewardbench2 · Paperjiang2026sosbench · Paperchaudhury2025chameleonbench · Papertan2025judgebench · Paperwhite2025livebench · Papersong2025prmbench · Paperragrewardbench2025 · Paperchen2025cotfaithfulness · Paperlambert2025rewardbench · Paperlin2025wildbench · Scholarchiang2024arena · Scholarguan2024hallusionbench · Scholarmazeika2024harmbench · Paperniu2024ragtruth · Paperzhang2024safetybench · Paperrottger2024xstest · Paperli2023alpacaeval · Papergao2023citations · Papermin2023factscore · Paperzhou2023ifeval · Paperzheng2023mtbench · Paperlin2022truthfulqa · Paperbrito2026safetyisnotuniversal · DOIhahm2026alignmenttampering · Paperchaudhari2026erts · Paperhung2026grayzone · Paperlian2026ethicalvulnerability · DOIwang2026safetymem · Paperchoi2026context · Paperlong2026languagegames · Papersharma2025constitutionalclassifiers · Paperbetley2025emergent · Papershen2025antidote · Paperandriushchenko2025adaptive · Paperwang2025lifelong · Papergreenblatt2025alignmentfaking · Paperren2024codeattack · DOIsouly2024strongreject · Paperwang2024backdooralign · Paperqi2024finetuning · Papermazeika2024harmbench · Paperzou2024circuitbreakers · Paperhubinger2024sleeper · Paperdenison2024subterfuge · Paperrottger2024xstest · Paperliang2023helm · Scholarshevlane2023extremerisks · Paperzou2023jailbreak · Paperlangosco2022goal · Papershah2022goal · Paperperez2022redteaming · DOIamodei2016concrete · Paperzhang2026safetyoneshot · Paperbrazilek2026alignment · Paperbach2026continual · Paperwang2026sot · Paperdjuhera2026safemerge · Paperhuang2025antidote · Paperzhang2025copr · Paperbetley2025emergent · Paperli2025larf · Paperwang2025lifelong · Paperli2025salora · Scholarzhao2025safetyneurons · Paperqi2024finetuning · Paperzou2024circuitbreakers · Paperlyu2024keeping · Paperhuang2024lisa · Paperhuang2024vaccine · Papernaseem2026mechanistic · arXivnayebi2026barriers · Paperlian2026ethicalvulnerability · DOIyong2026selfjailbreaking · Paperliu2026statisticalalignment · DOImazzu2026sustaining · Paperfalahati2026alignmentgame · Papersharkey2025openproblems · arXivyao2025alignmenttrap · arXivhudson2025corrigibility · Papergolz2025distortion · Paperchen2025cotfaithfulness · Paperwollschlager2025geometry · Paperpan2025hiddendimensions · Paperchen2025safetyneurons · Paperzhao2025safetyneurons · Papersheshadri2025alignmentfaking · DOIleong2025template · Paperlieberum2024gemmascope · arXivtempleton2024scaling · Paperwolf2024limitations · Paperarditi2024refusal · Paperrafailov2024overoptimization · Paperrimsky2024caa · DOIim2024dynamics · Paperjain2024safetyfinetuning · Paperburns2023latent · Scholarli2023iti · DOIzou2023repreng · Paperhadfieldmenell2017offswitch · DOIsoares2015corrigibility · Paperfeng2026benchmarking · arXivye2026codingenemy · arXivalamdari2026formal · DOIbaek2026sycophancy · arXivyu2026whensayingno · DOIwang2026agenticeval · Paperuluirmak2026evalsafetygap · Paperpanfilov2026dishonesty · Paperzhong2026weights · Paperwen2025adaptivedeployment · Paperzheng2025calm · DOIrichter2025auditing · Papermarks2025auditing · Paperarnav2025cot · DOImckenzie2025highstakes · DOIli2025streaming · DOIlambert2025rewardbench · Papersheshadri2025alignmentfaking · DOIli2025safetyanalyst · Papergreenblatt2025alignmentfaking · Paperamirizaniani2024auditllm · DOIgreenblatt2024control · Papercasper2024blackbox · DOIoren2024contamination · Scholarhubinger2024sleeper · Paperdenison2024subterfuge · Paperperez2023modelwritten · Paperliang2023helm · Scholarshevlane2023extremerisks · PaperThese surveys span several branches and are kept outside any single technical category.
ji2025alignment · Paperyu2025multimodalsurvey · Paperjiang2024preferencesurvey · Papergao2024unifiedpreference · Paperfernandes2024feedback · Papershen2023alignment · PaperContributions and bibliographic corrections are welcome:
Suggested inclusion criteria:
This is a curated and evolving research map rather than a claim of exhaustive coverage. Placement indicates relevance to a branch; it does not imply that a paper solves Human–AI Alignment or establishes a deployment guarantee.
Last synchronized with the taxonomy and bibliography on 2026-09-29. Citation aliases are normalized in the collection so that the same work is not counted twice under different BibTeX keys.
A taxonomy-guided collection of research on specifying, supervising, implementing, and assuring Human–AI Alignment.
This survey provides a unified overview of Human–AI Alignment. We propose a lifecycle taxonomy that organizes the field along four dimensions—alignment specification, supervision, mechanisms, and assurance—and use it to structure the literature across 23 non-exclusive research branches.
Reference components for 22 cited papers in Task & Assistance, Personalized Alignment, and Uncertainty & Drift are documented in Alignment specification, with explicit implementation limits and an offline demo.
Human Feedback and AI Feedback code, covering the supervision components of 17 cited papers, is documented in Alignment supervision.
The repository also provides a Python framework that mirrors the four lifecycle dimensions while keeping common workflows simple. The core has no runtime dependencies; training libraries are installed only for the methods that need them.
python -m pip install -e .
hai-align catalog validate
python examples/minimal_pipeline.py
To train with DPO:
python -m pip install -e ".[dpo]"
hai-align train dpo --model Qwen/Qwen3-0.6B --dataset trl-lib/ultrafeedback_binarized --output outputs/qwen-dpo
from human_alignment import DPO
run = DPO(
model="Qwen/Qwen3-0.6B",
dataset="trl-lib/ultrafeedback_binarized",
output_dir="outputs/qwen-dpo",
).train()
print(run.generate("What is Human--AI alignment?"))
See Codebase architecture for the public API and extension points. The machine-readable taxonomy lives in src/human_alignment/catalog/data/taxonomy.json.
Preference distillation is available through one consistent API for VPD, PPD, DCKD, TVKD, ADPA, and CTPD:
python -m pip install -e ".[distillation]"
from human_alignment import VPD, PreferenceDistillationExample
data = [
PreferenceDistillationExample(
prompt="Explain alignment briefly.",
responses=("Alignment connects behavior to human targets.", "It is model scaling."),
teacher_scores=(1.0, 0.0),
)
]
run = VPD(model="student-model", dataset=data, output_dir="outputs/student-vpd").train()
See Preference distillation for objective-specific dataset schemas, teacher-model use, and migration details.
The benchmarks the survey uses as evidence (RewardBench 2, JudgeBench, AlpacaEval 2) run under their official scoring rules, and three inference-time methods (CAA, best-of-N with a reward model, ARGS) steer a frozen model:
hai-align bench rewardbench2 --reward-model Skywork/Skywork-Reward-V2-Qwen3-0.6B --output-dir eval/rb2
See Benchmarks and inference-time methods.
Safety-alignment training is integrated under the same package, with 25 stages covering SafeRLHF, SafeDPO, BSO, SACPO, CAN, MODPO, CPO, BFPO, MidPO, reward and cost modeling, and multi-objective RLHF:
python -m pip install -e ".[safety]"
hai-align safety list
hai-align safety run recipes/safety_alignment/methods/safedpo.yaml
from human_alignment import SafetyAlignment
run = SafetyAlignment(
config="recipes/safety_alignment/methods/safedpo.yaml",
output_dir="outputs/safedpo",
).train()
See Safety alignment for recipe dependencies, paper-faithful implementation choices, evaluation, and H100 workflows.
Additional preference methods are available as IPO, BPO, TDPO, TISDPO,
TIDPO, TBPOQ, and TBPOA, with recipes under recipes/preference_optimization/.
See Preference optimization for installation,
token-weight data formats, objective sources, and checkpoint handling.
The library supports both ends of an alignment experiment:
from human_alignment import DPO, load_checkpoint
# Fine-tune a pretrained model or an existing checkpoint.
trained = DPO(
model="models/student_sft",
dataset="data/preferences",
output_dir="outputs/student_dpo",
).train()
# Load the resulting checkpoint later for inference or assurance.
checkpoint = load_checkpoint(
"outputs/student_dpo",
config={"device_map": "auto", "torch_dtype": "bfloat16"},
)
print(checkpoint.generate("Explain alignment."))
Preference-distillation supervision can be prepared without the temporary research scripts:
hai-align prepare distillation dckd \
--dataset HuggingFaceH4/ultrafeedback_binarized \
--teacher models/teacher_dpo \
--tokenizer models/student_sft \
--output data/ultrafeedback-dckd
| Example | Purpose |
|---|---|
minimal_pipeline.py | Validate the catalog and run a minimal end-to-end pipeline |
dpo_quickstart.py | Launch DPO from Python |
dpo.toml | Configure a DPO run declaratively |
checkpoint_workflow.py | Train, reload, and use a checkpoint |
preference_distillation.py | Prepare and run preference distillation |
Human–AI alignment is organized as a lifecycle spanning alignment specification, supervision, mechanisms, and assurance across chat, code, mathematical reasoning, multimodal, agentic, robotics, and healthcare settings.
View the high-resolution survey overview (PDF)
The lifecycle taxonomy organizes Human–AI alignment into four dimensions and eight groups; the paper collection below further resolves them into 23 non-exclusive research branches.
View the high-resolution taxonomy (PDF)
| Dimension | Guiding question | Groups |
|---|---|---|
| Alignment Specification | What should an AI system align to, and whose objectives and values should count? | Alignment Objectives · Values & Stakeholders |
| Alignment Supervision | Where do alignment signals come from, and how are they expressed and scaled? | Feedback Source · Feedback & Oversight |
| Alignment Mechanisms | How are alignment signals translated into model behavior during training and inference? | Training-Time Alignment · Inference-Time Alignment |
| Alignment Assurance | How do we evaluate, stress-test, preserve, interpret, and monitor alignment? | Evaluation & Robustness · System Assurance |
Training-time alignment covers reward and verifier modeling, supervised alignment, preference optimization, reinforcement learning, and alignment distillation.
View the high-resolution training-time diagram (PDF)
Inference-time alignment covers steering, search, iterative refinement, and interactive or agentic control while keeping the model policy fixed.
View the high-resolution inference-time diagram (PDF)
Feedback granularity and optimization granularity are tracked separately because the unit receiving feedback can differ from the unit optimized by the learning objective.
Papers are sorted by year within each branch. Each entry links to the publication page, DOI, arXiv record, or a clearly labeled Scholar search when the BibTeX record has no direct link.
What should an AI system align to, and whose objectives and values should count?
| Group | Branch | Papers |
|---|---|---|
| Alignment Objectives | Task & Assistance Alignment | 35 |
| Alignment Objectives | Safety Alignment | 27 |
| Values & Stakeholders | Personalized Alignment | 24 |
| Values & Stakeholders | Pluralistic & Societal Alignment | 35 |
| Values & Stakeholders | Context, Uncertainty & Drift | 23 |
purpura2026mosaic · DOIkwon2026reasonif · DOIdong2026ifevalpp · DOIrobinette2026verify · DOIxiao2026sftmix · Paperli2025fbbench · Paperzhang2025iopo · Paperzhang2025longreward · Paperharada2025manyifeval · Paperbalepur2025goodplan · Paperxu2025magpie · Paperan2025ultraif · Paperliu2024llavarlhf · Paperwang2024helpsteer2 · DOIyu2024lions · Papersun2024parrot · Paperniu2024ragtruth · Paperchung2024flan · Papercui2024ultrafeedback · Paperrafailov2023dpo · Papergao2023citations · Paperwu2023finegrained · Paperzhou2023lima · Paperwang2023selfinstruct · Paperdong2023steerlm · Paperlongpre2023flancollection · Papermishra2022natural · Paperwei2022flan · Paperglaese2022sparrow · Papersanh2022t0 · Paperwang2022supernatural · Paperbai2022hh · Paperouyang2022instructgpt · Paperaskell2021assistant · Papernakano2021webgpt · Paperzhang2026llmva · Paperjiang2026els · Paperwang2026hiar · Paperin2026alttrain · Paperbrito2026safetyisnotuniversal · DOIzhang2026alphaalign · Paperyang2026lasa · Paperniu2026nspo · Paperzhang2026safetyoneshot · Paperkim2026safedpo · Paperzhang2025controllablesafetyalignment · Paperkaraman2025porover · Paperli2025explicitsafety · Papergao2025abd · Paperzhang2025constrainedalignment · Paperhuang2025antidote · Paperli2025larf · Paperwang2025lifelong · Paperguan2024deliberative · Paperzou2024circuitbreakers · Papermu2024rulebased · Scholardai2024saferlhf · Paperhuang2024vaccine · Paperji2023beavertails · Paperbai2022constitutional · Paperbai2022hh · Paperaskell2021assistant · Papergarbacea2026personalizedbenchmark · Papersun2026fpps · Paperli2026alignx · Papertan2026profilepeft · Paperma2026cbpo · Paperwu2025interaction · Paperkim2025drift · Paperthonet2025fast · Paperbu2025cope · Papercho2025ticl · Paperbalepur2025personas · Paperguan2025personalizedsurvey · Paperaroca2025prose · Paperchen2025pal · Paperzhu2025personality · Paperpadurean2025useralign · Paperchen2025pad · Paperzhang2025personajudge · Paperzhao2025rlpa · Papertan2024oppu · Paperli2024personalized · Paperpoddar2024personalizing · Paperkirk2024prism · Paperjang2023personalizedsoups · Paperpooledayan2026overtonbench · Papertan2026subgroupvalues · Paperzhang2026cultivating · Papersun2026cuma · Paperzhou2026disalign · Papernie2026perspectra · Paperkim2026valueflow · Paperzhang2026crosscultural · Paperbhattacharyya2026alpha · DOIzhu2026dvmap · Paperpan2026fgdalign · DOIzhang2026culturemanager · Paperali2026pluralistic · DOIengelmann2025pluralisticcompatibility · Papermentxaka2025democracy · Papershirali2025heterogeneous · DOIzhang2025pluralisticcot · Paperki2025culturaldebate · Paperhalpern2025pairwise · Paperchen2025pal · Paperyao2025gdpo · Paperxu2025culturespa · Paperadams2025steerable · DOIgreenblatt2024control · Paperhuang2024collective · Paperli2024culturellm · Papersiththaranjan2024distributional · Scholarfeng2024modular · Papersorensen2024pluralistic · Scholarconitzer2024socialchoice · Papermu2024rulebased · Scholarkirk2024prism · Paperbai2022constitutional · Paperbakker2022finetuning · Paperawad2018moral · DOIfang2026aam · DOIlin2026activedpo · Paperchen2026cfa · DOIxu2026fadpo · Paperkeswani2026moralchange · Paperwu2026rmrouting · Papershen2025activerm · Paperfeng2025pilaf · Paperxu2025robustdpo · Paperli2025uipo · Paperxu2025uncertaintyjudge · Paperzhang2025copr · Paperwang2025lifelong · Papersun2025ugda · Papermuldrew2024active · Scholarsiththaranjan2024distributional · Scholarboerstler2024stability · Paperkong2024perpcorrect · Paperhadfieldmenell2017ird · Scholarhadfieldmenell2017offswitch · DOIhadfieldmenell2016cirl · Scholarabbeel2004apprenticeship · Scholarng2000irl · ScholarWhere do alignment signals come from, and how are they expressed and scaled?
| Group | Branch | Papers |
|---|---|---|
| Feedback Source | Human Feedback | 20 |
| Feedback Source | AI Feedback | 23 |
| Feedback Source | Programmatic & Verifiable Feedback | 22 |
| Feedback & Oversight | Demonstrations & Preferences | 24 |
| Feedback & Oversight | Critique, Process & Trajectory Feedback | 19 |
| Feedback & Oversight | Reliable & Scalable Oversight | 25 |
shi2026wildfeedback · Papershaikh2025ditto · Paperjung2025bco · Papermiranda2025hybrid · Paperzhang2025mmrlhf · Paperliu2024llavarlhf · Paperxu2024finegrained · Paperhou2024chatglmrlhf · Paperwang2024helpsteer2 · DOIyu2024rlhfv · Paperlloret2024alt · Paperwang2023helpsteer · arXivkopf2023openassistant · arXivwu2023finegrained · Paperglaese2022sparrow · Paperbai2022hh · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperziegler2019finetuning · Paperchristiano2017preferences · Paperyin2026sao · Papercook2026checklist · Paperxu2025magpie · Paperyu2025daif · Paperye2025conj · Paperxiong2025llavacritic · DOImcaleese2024critics · Paperliu2024dlma · Paperlee2024rlaif · Paperyuan2024selfrewarding · Scholarcui2024ultrafeedback · Paperli2024vlfeedback · Paperding2023ultrachat · arXivliu2023geval · arXivzhu2023judgelm · arXivzheng2023mtbench · Papermukherjee2023orca · arXivshinn2023reflexion · Scholarsun2023selfalignment · Scholarwang2023selfinstruct · Papermadaan2023selfrefine · Paperxu2023wizardlm · arXivbai2022constitutional · Papersu2026rewardbridge · Paperyuan2026k2v · Paperzhang2025openprm · Paperdeepseek2025r1 · Paperzhang2025genrm · Papersetlur2025rewarding · Papergehring2025rlef · Paperdong2025autoif · Paperlightman2024verify · Papersingh2024restem · Paperxin2024deepseekprover15 · DOIshao2024deepseekmath · Paperwang2024mathshepherd · Papermu2024rulebased · Scholarhosseini2024vstar · Paperchen2023codet · Paperyang2023leandojo · Paperliu2023rltf · Paperli2022coderl · Paperzelikman2022star · Paperchen2021humaneval · arXivhendrycks2021apps · arXivshi2026wildfeedback · Papermiranda2025hybrid · Papershaikh2025ditto · Papertan2025pugc · Paperjung2025bco · Paperxiao2025sweetspot · Paperli2025joint · Paperxu2024finegrained · Paperchung2024flan · Papermuldrew2024active · Scholarli2024demoreward · Paperwang2024helpsteer2 · DOIyu2024rlhfv · Papercui2024ultrafeedback · Paperwu2023finegrained · Paperzhou2023lima · Paperwang2023selfinstruct · Paperlongpre2023flancollection · Papermishra2022natural · Paperwei2022flan · Paperwang2022supernatural · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperchristiano2017preferences · Paperxi2026critiquerl · Paperfan2026agentrrm · Paperzhang2026biprm · Paperyin2025dgprm · Paperzhang2025openprm · Papersetlur2025rewarding · Paperyu2025criticrm · Paperxie2025ctrl · Paperyu2025rco · Papergou2024critic · Scholarlightman2024verify · Papermcaleese2024critics · Paperyu2024rlhfv · Papercui2024ultrafeedback · Paperwu2023finegrained · Papershinn2023reflexion · Scholarmadaan2023selfrefine · Paperbai2022constitutional · Paperuesato2022process · Paperranaldi2026agentic · Paperlang2026selective · DOIyin2026partitioned · Paperqiu2026peerprediction · Paperrahman2025aidebate · Scholarnayebi2025barriers · Paperlang2025debate · Papergoel2025oversight · Paperhammond2025neuralproofs · Paperlambert2025rewardbench · Paperengels2025scaling · Papermuldrew2024active · Scholargreenblatt2024control · Paperkhan2024debate · Papersun2024easyhard · Paperkenton2024scalableoversight · Paperkirchner2024proververifier · Paperwang2024rmquality · DOIcheng2024rime · Scholarzhang2024selm · Paperburns2024weakstrong · Paperwu2023finegrained · Papergao2023overoptimization · Paperirving2018debate · Paperchristiano2018amplification · PaperHow are alignment signals translated into model behavior during training and inference?
| Group | Branch | Papers |
|---|---|---|
| Training-Time Alignment | Reward & Verifier Modeling | 27 |
| Training-Time Alignment | Supervised Alignment | 22 |
| Training-Time Alignment | Preference Optimization | 50 |
| Training-Time Alignment | Reinforcement Learning | 32 |
| Training-Time Alignment | Alignment Distillation | 9 |
| Inference-Time Alignment | Steering, Search & Refinement | 31 |
| Inference-Time Alignment | Interactive & Agentic Control | 17 |
liu2026openrubrics · Paperwang2026rationale · Paperzhou2026prismrm · Paperchen2026rmr1 · Paperzhang2026biprm · Papermiao2026adajudge · Paperpeng2025agenticrm · Paperyin2025dgprm · Paperzhang2025genrm · Paperwang2025gram · Paperxiong2025llavacritic · DOIguo2025rrm · DOIchen2025pal · Paperwang2024helpsteer2 · DOIlightman2024verify · Paperwang2024mathshepherd · Paperkirchner2024proververifier · Papercoste2024ensembles · Paperdai2024saferlhf · Paperwu2023finegrained · Paperglaese2022sparrow · Paperuesato2022process · Paperbai2022hh · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperziegler2019finetuning · Paperchristiano2017preferences · Paperlee2026oasis · Paperxiao2026sftmix · Paperxu2025magpie · Paperyang2025main · Paperharada-etal-2025-massive · Paperhua-etal-2025-intuitive · Paperzhang2025grape · Scholarxia2024less · Paperyang2024metaaligner · DOIchung2024flan · Paperliu2024selectit · Paperli2024backtranslation · Paperliu2024deita · Paperzhou2023lima · Paperwang2023selfinstruct · Paperdong2023steerlm · Paperlongpre2023flancollection · Papermishra2022natural · Paperwei2022flan · Papersanh2022t0 · Paperwang2022supernatural · Paperouyang2022instructgpt · Paperle2026causaldpo · Paperzhu2026rappo · Paperyuan2026hypo · Paperliu2026statisticalalignment · DOIoi2026autoregressive · Papernguyen2026bso · DOIkim2026safedpo · Paperluo2026sharpness · Paperyang2026token · Papernguyen2026tokenratio · Paperzhang2025constrainedalignment · Paperwu2025alphadpo · Paperdoosterlinck2025apo · Paperxu2025doublyrobust · Paperlee2025kl · Paperwang2025mpo · Paperzhu2025mmedpo · Paperyao2025gdpo · Paperkim2025bregmanpo · Paperguo2025pro · Papersun2025rpo · Paperzhang2025risk · Paperkong2025sdpo · Paperwu2025sppo · Paperzhu2025sgdpo · Paperzhou2025treg · Paperliu2025tisdpo · Paperwu2025drdpo · Paperzhu2025wspo · Papermuldrew2024active · Scholarxu2024behaviorbpo · Paperxu2024cpo · Paperrosset2024dno · Paperethayarajh2024kto · Paperyang2024metaaligner · DOIshani2024multiturn · Scholarmunos2024nash · Paperhong2024orpo · Paperchowdhury2024rdpo · Papercheng2024rime · Scholarding2024sail · Paperzhang2024selm · Papermeng2024simpo · Paperzeng2024tdpo · Paperwu2024beta · Paperazar2023psipo · Paperzhao2023hadpo · Paperrafailov2023dpo · Paperyuan2023rrhf · Paperzhao2023slichf · Paperzhang2026alphaalign · Paperjiang2026rlvrr · Paperniu2026nspo · Paperzhang2026corewarding · Paperchu2026gpg · Paperzeng2026rlve · arXivdeng2026srl · Paperviswanathan2025rlcf · Paperliu2025dapo_advantage · Paperdeepseek2025r1 · Paperpeng2025repo · Paperzhang2025mmrlhf · Paperwalder2025passk · Paperliu2025saturn · Paperguo2025spo · Paperfang2025serl · Paperleroux2025tapered · Paperliu2024llavarlhf · Paperahmadian2024reinforce · Paperhou2024chatglmrlhf · Papershao2024deepseekmath · Papershani2024multiturn · Scholarli2024remax · Paperlee2024rlaif · Papermu2024rulebased · Scholardai2024saferlhf · Paperding2024sail · Paperkaufmann2023survey · Paperuesato2022process · Paperouyang2022instructgpt · Paperstiennon2020summarize · Paperchristiano2017preferences · Papernguyen2026ctpd · Papergao2025adpa · Paperzhang2025aligndistil · Papergu2025pad · Paperkwon2025tvkd · Paperhong2024cyclealign · Paperliu2024dlma · Paperli2024dpkd · Paperzhang2024plad · Paperli2026trainingfree · arXivnguyen2026beyondlinear · arXivwu2026cybercorrect · arXivluo2026dontlosefocus · arXivzhao2026odesteer · arXivyou2026spherical · arXivnguyen2026pidsteering · Paperweng2026finesteer · Paperzhou2026requer · Papersnell2025scaling · Papershi2025ssr · arXivli2025unleashing · arXivding2025w2s · arXivmasters2025arcane · Paperzhang2025controllablesafetyalignment · Paperstolfo2025steering · Paperxiong2025llavacritic · DOIfei2025nudging · Paperhung2025darwin · Paperma2025s2r · Papermudgal2023controlled · Paperkhanov2024args · Papergou2024critic · Scholarwang2024mathshepherd · Paperrimsky2024caa · DOIturner2023activationaddition · Paperli2023iti · DOIshinn2023reflexion · Scholarzou2023repreng · Papermadaan2023selfrefine · Papernakano2021webgpt · Paperzhou2026steering · arXivzhu2026oversight · Papersuri2026clarification · Paperayankoya2025scalable · DOIshirali2025burden · arXivhudson2025corrigibility · Paperhe2025aria · Paperxue2025agsa · Papersouth2025delegation · Papergarber2025offswitch · Papergreenblatt2024control · Paperandukuri2024stargate · Scholarmozannar2020defer · Scholarhadfieldmenell2017offswitch · DOIhadfieldmenell2016cirl · Scholarorseau2016interruptible · Scholarsoares2015corrigibility · PaperHow do we evaluate, stress-test, preserve, interpret, and monitor alignment?
| Group | Branch | Papers |
|---|---|---|
| Evaluation & Robustness | Behavioral & Evaluator Evaluation | 26 |
| Evaluation & Robustness | Adversarial & Distribution Robustness | 30 |
| Evaluation & Robustness | Alignment Preservation | 17 |
| System Assurance | Mechanistic & Theoretical Evidence | 31 |
| System Assurance | Monitoring & Auditing | 29 |
wang2026planrewardbench · DOIwen2026ifrewardbench · Paperlaban2026lost · Paperxiong2026multicrit · Papermalik2026rewardbench2 · Paperjiang2026sosbench · Paperchaudhury2025chameleonbench · Papertan2025judgebench · Paperwhite2025livebench · Papersong2025prmbench · Paperragrewardbench2025 · Paperchen2025cotfaithfulness · Paperlambert2025rewardbench · Paperlin2025wildbench · Scholarchiang2024arena · Scholarguan2024hallusionbench · Scholarmazeika2024harmbench · Paperniu2024ragtruth · Paperzhang2024safetybench · Paperrottger2024xstest · Paperli2023alpacaeval · Papergao2023citations · Papermin2023factscore · Paperzhou2023ifeval · Paperzheng2023mtbench · Paperlin2022truthfulqa · Paperbrito2026safetyisnotuniversal · DOIhahm2026alignmenttampering · Paperchaudhari2026erts · Paperhung2026grayzone · Paperlian2026ethicalvulnerability · DOIwang2026safetymem · Paperchoi2026context · Paperlong2026languagegames · Papersharma2025constitutionalclassifiers · Paperbetley2025emergent · Papershen2025antidote · Paperandriushchenko2025adaptive · Paperwang2025lifelong · Papergreenblatt2025alignmentfaking · Paperren2024codeattack · DOIsouly2024strongreject · Paperwang2024backdooralign · Paperqi2024finetuning · Papermazeika2024harmbench · Paperzou2024circuitbreakers · Paperhubinger2024sleeper · Paperdenison2024subterfuge · Paperrottger2024xstest · Paperliang2023helm · Scholarshevlane2023extremerisks · Paperzou2023jailbreak · Paperlangosco2022goal · Papershah2022goal · Paperperez2022redteaming · DOIamodei2016concrete · Paperzhang2026safetyoneshot · Paperbrazilek2026alignment · Paperbach2026continual · Paperwang2026sot · Paperdjuhera2026safemerge · Paperhuang2025antidote · Paperzhang2025copr · Paperbetley2025emergent · Paperli2025larf · Paperwang2025lifelong · Paperli2025salora · Scholarzhao2025safetyneurons · Paperqi2024finetuning · Paperzou2024circuitbreakers · Paperlyu2024keeping · Paperhuang2024lisa · Paperhuang2024vaccine · Papernaseem2026mechanistic · arXivnayebi2026barriers · Paperlian2026ethicalvulnerability · DOIyong2026selfjailbreaking · Paperliu2026statisticalalignment · DOImazzu2026sustaining · Paperfalahati2026alignmentgame · Papersharkey2025openproblems · arXivyao2025alignmenttrap · arXivhudson2025corrigibility · Papergolz2025distortion · Paperchen2025cotfaithfulness · Paperwollschlager2025geometry · Paperpan2025hiddendimensions · Paperchen2025safetyneurons · Paperzhao2025safetyneurons · Papersheshadri2025alignmentfaking · DOIleong2025template · Paperlieberum2024gemmascope · arXivtempleton2024scaling · Paperwolf2024limitations · Paperarditi2024refusal · Paperrafailov2024overoptimization · Paperrimsky2024caa · DOIim2024dynamics · Paperjain2024safetyfinetuning · Paperburns2023latent · Scholarli2023iti · DOIzou2023repreng · Paperhadfieldmenell2017offswitch · DOIsoares2015corrigibility · Paperfeng2026benchmarking · arXivye2026codingenemy · arXivalamdari2026formal · DOIbaek2026sycophancy · arXivyu2026whensayingno · DOIwang2026agenticeval · Paperuluirmak2026evalsafetygap · Paperpanfilov2026dishonesty · Paperzhong2026weights · Paperwen2025adaptivedeployment · Paperzheng2025calm · DOIrichter2025auditing · Papermarks2025auditing · Paperarnav2025cot · DOImckenzie2025highstakes · DOIli2025streaming · DOIlambert2025rewardbench · Papersheshadri2025alignmentfaking · DOIli2025safetyanalyst · Papergreenblatt2025alignmentfaking · Paperamirizaniani2024auditllm · DOIgreenblatt2024control · Papercasper2024blackbox · DOIoren2024contamination · Scholarhubinger2024sleeper · Paperdenison2024subterfuge · Paperperez2023modelwritten · Paperliang2023helm · Scholarshevlane2023extremerisks · PaperThese surveys span several branches and are kept outside any single technical category.
ji2025alignment · Paperyu2025multimodalsurvey · Paperjiang2024preferencesurvey · Papergao2024unifiedpreference · Paperfernandes2024feedback · Papershen2023alignment · PaperContributions and bibliographic corrections are welcome:
Suggested inclusion criteria:
This is a curated and evolving research map rather than a claim of exhaustive coverage. Placement indicates relevance to a branch; it does not imply that a paper solves Human–AI Alignment or establishes a deployment guarantee.
Last synchronized with the taxonomy and bibliography on 2026-09-29. Citation aliases are normalized in the collection so that the same work is not counted twice under different BibTeX keys.