Thang1703hrsh/Awesome-Human-AI-Alignment

Python

96

30 commits

updated Oct 3, 2026

See the code

README

Awesome Human–AI Alignment

A taxonomy-guided collection of research on specifying, supervising, implementing, and assuring Human–AI Alignment.

Dimensions Terminal branches Works PRs

News

Overview

This survey provides a unified overview of Human–AI Alignment. We propose a lifecycle taxonomy that organizes the field along four dimensions—alignment specification, supervision, mechanisms, and assurance—and use it to structure the literature across 23 non-exclusive research branches.

Software framework

Reference components for 22 cited papers in Task & Assistance, Personalized Alignment, and Uncertainty & Drift are documented in Alignment specification, with explicit implementation limits and an offline demo.

Human Feedback and AI Feedback code, covering the supervision components of 17 cited papers, is documented in Alignment supervision.

The repository also provides a Python framework that mirrors the four lifecycle dimensions while keeping common workflows simple. The core has no runtime dependencies; training libraries are installed only for the methods that need them.

python -m pip install -e .
hai-align catalog validate
python examples/minimal_pipeline.py

To train with DPO:

python -m pip install -e ".[dpo]"
hai-align train dpo --model Qwen/Qwen3-0.6B --dataset trl-lib/ultrafeedback_binarized --output outputs/qwen-dpo
from human_alignment import DPO

run = DPO(
    model="Qwen/Qwen3-0.6B",
    dataset="trl-lib/ultrafeedback_binarized",
    output_dir="outputs/qwen-dpo",
).train()

print(run.generate("What is Human--AI alignment?"))

See Codebase architecture for the public API and extension points. The machine-readable taxonomy lives in src/human_alignment/catalog/data/taxonomy.json.

Preference distillation is available through one consistent API for VPD, PPD, DCKD, TVKD, ADPA, and CTPD:

python -m pip install -e ".[distillation]"
from human_alignment import VPD, PreferenceDistillationExample

data = [
    PreferenceDistillationExample(
        prompt="Explain alignment briefly.",
        responses=("Alignment connects behavior to human targets.", "It is model scaling."),
        teacher_scores=(1.0, 0.0),
    )
]
run = VPD(model="student-model", dataset=data, output_dir="outputs/student-vpd").train()

See Preference distillation for objective-specific dataset schemas, teacher-model use, and migration details.

The benchmarks the survey uses as evidence (RewardBench 2, JudgeBench, AlpacaEval 2) run under their official scoring rules, and three inference-time methods (CAA, best-of-N with a reward model, ARGS) steer a frozen model:

hai-align bench rewardbench2 --reward-model Skywork/Skywork-Reward-V2-Qwen3-0.6B --output-dir eval/rb2

See Benchmarks and inference-time methods.

Safety-alignment training is integrated under the same package, with 25 stages covering SafeRLHF, SafeDPO, BSO, SACPO, CAN, MODPO, CPO, BFPO, MidPO, reward and cost modeling, and multi-objective RLHF:

python -m pip install -e ".[safety]"
hai-align safety list
hai-align safety run recipes/safety_alignment/methods/safedpo.yaml
from human_alignment import SafetyAlignment

run = SafetyAlignment(
    config="recipes/safety_alignment/methods/safedpo.yaml",
    output_dir="outputs/safedpo",
).train()

See Safety alignment for recipe dependencies, paper-faithful implementation choices, evaluation, and H100 workflows.

Additional preference methods are available as IPO, BPO, TDPO, TISDPO, TIDPO, TBPOQ, and TBPOA, with recipes under recipes/preference_optimization/. See Preference optimization for installation, token-weight data formats, objective sources, and checkpoint handling.

The library supports both ends of an alignment experiment:

from human_alignment import DPO, load_checkpoint

# Fine-tune a pretrained model or an existing checkpoint.
trained = DPO(
    model="models/student_sft",
    dataset="data/preferences",
    output_dir="outputs/student_dpo",
).train()

# Load the resulting checkpoint later for inference or assurance.
checkpoint = load_checkpoint(
    "outputs/student_dpo",
    config={"device_map": "auto", "torch_dtype": "bfloat16"},
)
print(checkpoint.generate("Explain alignment."))

Preference-distillation supervision can be prepared without the temporary research scripts:

hai-align prepare distillation dckd \
  --dataset HuggingFaceH4/ultrafeedback_binarized \
  --teacher models/teacher_dpo \
  --tokenizer models/student_sft \
  --output data/ultrafeedback-dckd

Examples

ExamplePurpose
minimal_pipeline.pyValidate the catalog and run a minimal end-to-end pipeline
dpo_quickstart.pyLaunch DPO from Python
dpo.tomlConfigure a DPO run declaratively
checkpoint_workflow.pyTrain, reload, and use a checkpoint
preference_distillation.pyPrepare and run preference distillation

Taxonomy

Survey overview

Overview of the Human–AI Alignment survey framework and application settings

Human–AI alignment is organized as a lifecycle spanning alignment specification, supervision, mechanisms, and assurance across chat, code, mathematical reasoning, multimodal, agentic, robotics, and healthcare settings.

View the high-resolution survey overview (PDF)

Lifecycle taxonomy

Human–AI Alignment taxonomy

The lifecycle taxonomy organizes Human–AI alignment into four dimensions and eight groups; the paper collection below further resolves them into 23 non-exclusive research branches.

View the high-resolution taxonomy (PDF)

DimensionGuiding questionGroups
Alignment SpecificationWhat should an AI system align to, and whose objectives and values should count?Alignment Objectives · Values & Stakeholders
Alignment SupervisionWhere do alignment signals come from, and how are they expressed and scaled?Feedback Source · Feedback & Oversight
Alignment MechanismsHow are alignment signals translated into model behavior during training and inference?Training-Time Alignment · Inference-Time Alignment
Alignment AssuranceHow do we evaluate, stress-test, preserve, interpret, and monitor alignment?Evaluation & Robustness · System Assurance

Training-time mechanisms at a glance

Five training-time alignment mechanisms

Training-time alignment covers reward and verifier modeling, supervised alignment, preference optimization, reinforcement learning, and alignment distillation.

View the high-resolution training-time diagram (PDF)

Inference-time mechanisms at a glance

Four inference-time alignment mechanisms

Inference-time alignment covers steering, search, iterative refinement, and interactive or agentic control while keeping the model policy fixed.

View the high-resolution inference-time diagram (PDF)

Collection coverage

  • 172 representative assignments covering 167 citation keys in the compact figure, before alias normalization.
  • 598 collection assignments across 23 terminal branches; repeated placement is intentional.
  • 451 unique normalized works are linked from the collection.
  • 6 cross-cutting surveys are listed separately after the branch collection.

Cross-cutting descriptors

  • System and modality: text, multimodal, agentic, and embodied systems.
  • Feedback granularity: response, span, reasoning step, token, and action.
  • Optimization granularity: sequence, segment, token, and trajectory.
  • Learning regime: offline, online, static, and continual learning.

Feedback granularity and optimization granularity are tracked separately because the unit receiving feedback can differ from the unit optimized by the learning objective.

Paper collection

Papers are sorted by year within each branch. Each entry links to the publication page, DOI, arXiv record, or a clearly labeled Scholar search when the BibTeX record has no direct link.

Alignment Specification

What should an AI system align to, and whose objectives and values should count?

GroupBranchPapers
Alignment ObjectivesTask & Assistance Alignment35
Alignment ObjectivesSafety Alignment27
Values & StakeholdersPersonalized Alignment24
Values & StakeholdersPluralistic & Societal Alignment35
Values & StakeholdersContext, Uncertainty & Drift23

Alignment Objectives

Task & Assistance Alignment (35)

Safety Alignment (27)

Values & Stakeholders

Personalized Alignment (24)

Pluralistic & Societal Alignment (35)

Context, Uncertainty & Drift (23)

Alignment Supervision

Where do alignment signals come from, and how are they expressed and scaled?

GroupBranchPapers
Feedback SourceHuman Feedback20
Feedback SourceAI Feedback23
Feedback SourceProgrammatic & Verifiable Feedback22
Feedback & OversightDemonstrations & Preferences24
Feedback & OversightCritique, Process & Trajectory Feedback19
Feedback & OversightReliable & Scalable Oversight25

Feedback Source

Human Feedback (20)

AI Feedback (23)

Programmatic & Verifiable Feedback (22)

Feedback & Oversight

Demonstrations & Preferences (24)

Critique, Process & Trajectory Feedback (19)

Reliable & Scalable Oversight (25)

Alignment Mechanisms

How are alignment signals translated into model behavior during training and inference?

GroupBranchPapers
Training-Time AlignmentReward & Verifier Modeling27
Training-Time AlignmentSupervised Alignment22
Training-Time AlignmentPreference Optimization50
Training-Time AlignmentReinforcement Learning32
Training-Time AlignmentAlignment Distillation9
Inference-Time AlignmentSteering, Search & Refinement31
Inference-Time AlignmentInteractive & Agentic Control17

Training-Time Alignment

Reward & Verifier Modeling (27)

Supervised Alignment (22)

Preference Optimization (50)

Reinforcement Learning (32)

Alignment Distillation (9)

Inference-Time Alignment

Steering, Search & Refinement (31)

Interactive & Agentic Control (17)

Alignment Assurance

How do we evaluate, stress-test, preserve, interpret, and monitor alignment?

GroupBranchPapers
Evaluation & RobustnessBehavioral & Evaluator Evaluation26
Evaluation & RobustnessAdversarial & Distribution Robustness30
Evaluation & RobustnessAlignment Preservation17
System AssuranceMechanistic & Theoretical Evidence31
System AssuranceMonitoring & Auditing29

Evaluation & Robustness

Behavioral & Evaluator Evaluation (26)

Adversarial & Distribution Robustness (30)

Alignment Preservation (17)

System Assurance

Mechanistic & Theoretical Evidence (31)

Monitoring & Auditing (29)

General Survey Context

These surveys span several branches and are kept outside any single technical category.

Contributing

Contributions and bibliographic corrections are welcome:

  1. Include the paper's complete bibliographic record (authors, title, venue, year, and DOI or arXiv ID) in the pull request. Prefer the final venue page and DOI; use arXiv when no archival version is available.
  2. Add papers to every branch they substantively address; assignments do not need to be exclusive.
  3. Explain the proposed placement and include a stable public paper link in the pull request.

Suggested inclusion criteria:

  • The work contributes to alignment specification, supervision, training, inference-time control, evaluation, robustness, interpretability, or monitoring.
  • A stable manuscript or archival publication page is publicly available.
  • The bibliographic metadata can be independently checked.

Scope

This is a curated and evolving research map rather than a claim of exhaustive coverage. Placement indicates relevance to a branch; it does not imply that a paper solves Human–AI Alignment or establishes a deployment guarantee.


Last synchronized with the taxonomy and bibliography on 2026-09-29. Citation aliases are normalized in the collection so that the same work is not counted twice under different BibTeX keys.

Thang1703hrsh/Awesome-Human-AI-Alignment

Python

96

30 commits

updated Oct 3, 2026

See the code

README

Awesome Human–AI Alignment

A taxonomy-guided collection of research on specifying, supervising, implementing, and assuring Human–AI Alignment.

Dimensions Terminal branches Works PRs

News

Overview

This survey provides a unified overview of Human–AI Alignment. We propose a lifecycle taxonomy that organizes the field along four dimensions—alignment specification, supervision, mechanisms, and assurance—and use it to structure the literature across 23 non-exclusive research branches.

Software framework

Reference components for 22 cited papers in Task & Assistance, Personalized Alignment, and Uncertainty & Drift are documented in Alignment specification, with explicit implementation limits and an offline demo.

Human Feedback and AI Feedback code, covering the supervision components of 17 cited papers, is documented in Alignment supervision.

The repository also provides a Python framework that mirrors the four lifecycle dimensions while keeping common workflows simple. The core has no runtime dependencies; training libraries are installed only for the methods that need them.

python -m pip install -e .
hai-align catalog validate
python examples/minimal_pipeline.py

To train with DPO:

python -m pip install -e ".[dpo]"
hai-align train dpo --model Qwen/Qwen3-0.6B --dataset trl-lib/ultrafeedback_binarized --output outputs/qwen-dpo
from human_alignment import DPO

run = DPO(
    model="Qwen/Qwen3-0.6B",
    dataset="trl-lib/ultrafeedback_binarized",
    output_dir="outputs/qwen-dpo",
).train()

print(run.generate("What is Human--AI alignment?"))

See Codebase architecture for the public API and extension points. The machine-readable taxonomy lives in src/human_alignment/catalog/data/taxonomy.json.

Preference distillation is available through one consistent API for VPD, PPD, DCKD, TVKD, ADPA, and CTPD:

python -m pip install -e ".[distillation]"
from human_alignment import VPD, PreferenceDistillationExample

data = [
    PreferenceDistillationExample(
        prompt="Explain alignment briefly.",
        responses=("Alignment connects behavior to human targets.", "It is model scaling."),
        teacher_scores=(1.0, 0.0),
    )
]
run = VPD(model="student-model", dataset=data, output_dir="outputs/student-vpd").train()

See Preference distillation for objective-specific dataset schemas, teacher-model use, and migration details.

The benchmarks the survey uses as evidence (RewardBench 2, JudgeBench, AlpacaEval 2) run under their official scoring rules, and three inference-time methods (CAA, best-of-N with a reward model, ARGS) steer a frozen model:

hai-align bench rewardbench2 --reward-model Skywork/Skywork-Reward-V2-Qwen3-0.6B --output-dir eval/rb2

See Benchmarks and inference-time methods.

Safety-alignment training is integrated under the same package, with 25 stages covering SafeRLHF, SafeDPO, BSO, SACPO, CAN, MODPO, CPO, BFPO, MidPO, reward and cost modeling, and multi-objective RLHF:

python -m pip install -e ".[safety]"
hai-align safety list
hai-align safety run recipes/safety_alignment/methods/safedpo.yaml
from human_alignment import SafetyAlignment

run = SafetyAlignment(
    config="recipes/safety_alignment/methods/safedpo.yaml",
    output_dir="outputs/safedpo",
).train()

See Safety alignment for recipe dependencies, paper-faithful implementation choices, evaluation, and H100 workflows.

Additional preference methods are available as IPO, BPO, TDPO, TISDPO, TIDPO, TBPOQ, and TBPOA, with recipes under recipes/preference_optimization/. See Preference optimization for installation, token-weight data formats, objective sources, and checkpoint handling.

The library supports both ends of an alignment experiment:

from human_alignment import DPO, load_checkpoint

# Fine-tune a pretrained model or an existing checkpoint.
trained = DPO(
    model="models/student_sft",
    dataset="data/preferences",
    output_dir="outputs/student_dpo",
).train()

# Load the resulting checkpoint later for inference or assurance.
checkpoint = load_checkpoint(
    "outputs/student_dpo",
    config={"device_map": "auto", "torch_dtype": "bfloat16"},
)
print(checkpoint.generate("Explain alignment."))

Preference-distillation supervision can be prepared without the temporary research scripts:

hai-align prepare distillation dckd \
  --dataset HuggingFaceH4/ultrafeedback_binarized \
  --teacher models/teacher_dpo \
  --tokenizer models/student_sft \
  --output data/ultrafeedback-dckd

Examples

ExamplePurpose
minimal_pipeline.pyValidate the catalog and run a minimal end-to-end pipeline
dpo_quickstart.pyLaunch DPO from Python
dpo.tomlConfigure a DPO run declaratively
checkpoint_workflow.pyTrain, reload, and use a checkpoint
preference_distillation.pyPrepare and run preference distillation

Taxonomy

Survey overview

Overview of the Human–AI Alignment survey framework and application settings

Human–AI alignment is organized as a lifecycle spanning alignment specification, supervision, mechanisms, and assurance across chat, code, mathematical reasoning, multimodal, agentic, robotics, and healthcare settings.

View the high-resolution survey overview (PDF)

Lifecycle taxonomy

Human–AI Alignment taxonomy

The lifecycle taxonomy organizes Human–AI alignment into four dimensions and eight groups; the paper collection below further resolves them into 23 non-exclusive research branches.

View the high-resolution taxonomy (PDF)

DimensionGuiding questionGroups
Alignment SpecificationWhat should an AI system align to, and whose objectives and values should count?Alignment Objectives · Values & Stakeholders
Alignment SupervisionWhere do alignment signals come from, and how are they expressed and scaled?Feedback Source · Feedback & Oversight
Alignment MechanismsHow are alignment signals translated into model behavior during training and inference?Training-Time Alignment · Inference-Time Alignment
Alignment AssuranceHow do we evaluate, stress-test, preserve, interpret, and monitor alignment?Evaluation & Robustness · System Assurance

Training-time mechanisms at a glance

Five training-time alignment mechanisms

Training-time alignment covers reward and verifier modeling, supervised alignment, preference optimization, reinforcement learning, and alignment distillation.

View the high-resolution training-time diagram (PDF)

Inference-time mechanisms at a glance

Four inference-time alignment mechanisms

Inference-time alignment covers steering, search, iterative refinement, and interactive or agentic control while keeping the model policy fixed.

View the high-resolution inference-time diagram (PDF)

Collection coverage

  • 172 representative assignments covering 167 citation keys in the compact figure, before alias normalization.
  • 598 collection assignments across 23 terminal branches; repeated placement is intentional.
  • 451 unique normalized works are linked from the collection.
  • 6 cross-cutting surveys are listed separately after the branch collection.

Cross-cutting descriptors

  • System and modality: text, multimodal, agentic, and embodied systems.
  • Feedback granularity: response, span, reasoning step, token, and action.
  • Optimization granularity: sequence, segment, token, and trajectory.
  • Learning regime: offline, online, static, and continual learning.

Feedback granularity and optimization granularity are tracked separately because the unit receiving feedback can differ from the unit optimized by the learning objective.

Paper collection

Papers are sorted by year within each branch. Each entry links to the publication page, DOI, arXiv record, or a clearly labeled Scholar search when the BibTeX record has no direct link.

Alignment Specification

What should an AI system align to, and whose objectives and values should count?

GroupBranchPapers
Alignment ObjectivesTask & Assistance Alignment35
Alignment ObjectivesSafety Alignment27
Values & StakeholdersPersonalized Alignment24
Values & StakeholdersPluralistic & Societal Alignment35
Values & StakeholdersContext, Uncertainty & Drift23

Alignment Objectives

Task & Assistance Alignment (35)

Safety Alignment (27)

Values & Stakeholders

Personalized Alignment (24)

Pluralistic & Societal Alignment (35)

Context, Uncertainty & Drift (23)

Alignment Supervision

Where do alignment signals come from, and how are they expressed and scaled?

GroupBranchPapers
Feedback SourceHuman Feedback20
Feedback SourceAI Feedback23
Feedback SourceProgrammatic & Verifiable Feedback22
Feedback & OversightDemonstrations & Preferences24
Feedback & OversightCritique, Process & Trajectory Feedback19
Feedback & OversightReliable & Scalable Oversight25

Feedback Source

Human Feedback (20)

AI Feedback (23)

Programmatic & Verifiable Feedback (22)

Feedback & Oversight

Demonstrations & Preferences (24)

Critique, Process & Trajectory Feedback (19)

Reliable & Scalable Oversight (25)

Alignment Mechanisms

How are alignment signals translated into model behavior during training and inference?

GroupBranchPapers
Training-Time AlignmentReward & Verifier Modeling27
Training-Time AlignmentSupervised Alignment22
Training-Time AlignmentPreference Optimization50
Training-Time AlignmentReinforcement Learning32
Training-Time AlignmentAlignment Distillation9
Inference-Time AlignmentSteering, Search & Refinement31
Inference-Time AlignmentInteractive & Agentic Control17

Training-Time Alignment

Reward & Verifier Modeling (27)

Supervised Alignment (22)

Preference Optimization (50)

Reinforcement Learning (32)

Alignment Distillation (9)

Inference-Time Alignment

Steering, Search & Refinement (31)

Interactive & Agentic Control (17)

Alignment Assurance

How do we evaluate, stress-test, preserve, interpret, and monitor alignment?

GroupBranchPapers
Evaluation & RobustnessBehavioral & Evaluator Evaluation26
Evaluation & RobustnessAdversarial & Distribution Robustness30
Evaluation & RobustnessAlignment Preservation17
System AssuranceMechanistic & Theoretical Evidence31
System AssuranceMonitoring & Auditing29

Evaluation & Robustness

Behavioral & Evaluator Evaluation (26)

Adversarial & Distribution Robustness (30)

Alignment Preservation (17)

System Assurance

Mechanistic & Theoretical Evidence (31)

Monitoring & Auditing (29)

General Survey Context

These surveys span several branches and are kept outside any single technical category.

Contributing

Contributions and bibliographic corrections are welcome:

  1. Include the paper's complete bibliographic record (authors, title, venue, year, and DOI or arXiv ID) in the pull request. Prefer the final venue page and DOI; use arXiv when no archival version is available.
  2. Add papers to every branch they substantively address; assignments do not need to be exclusive.
  3. Explain the proposed placement and include a stable public paper link in the pull request.

Suggested inclusion criteria:

  • The work contributes to alignment specification, supervision, training, inference-time control, evaluation, robustness, interpretability, or monitoring.
  • A stable manuscript or archival publication page is publicly available.
  • The bibliographic metadata can be independently checked.

Scope

This is a curated and evolving research map rather than a claim of exhaustive coverage. Placement indicates relevance to a branch; it does not imply that a paper solves Human–AI Alignment or establishes a deployment guarantee.


Last synchronized with the taxonomy and bibliography on 2026-09-29. Citation aliases are normalized in the collection so that the same work is not counted twice under different BibTeX keys.