Skip to content

Latest commit

 

History

34 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AlignmentLab

AlignmentLab is a matched-compute comparison of SFT, DPO, and GRPO/RLVR post-training on Qwen3-8B and Qwen3-14B. The core question is whether reinforcement learning from verifiable rewards (RLVR) actually adds capability or only sharpens the sampling distribution — measured as pass@1 versus pass@k (k up to 256) on math reasoning. Every method is run under the same GPU-hour budget so that gains are attributable to the learning signal rather than to extra compute. Training uses TRL (SFT, DPO) and OpenRLHF (GRPO with Ray + vLLM); evaluation uses lm-evaluation-harness for standard benchmarks and math-verify for verifiable reward and pass@k scoring. Experiments run on A100/H100 nodes managed by IBM LSF (bsub/bjobs); no compute is ever run on the login node. Runs are logged to Weights & Biases (public project alignmentlab).

Research questions

  • RQ1: SFT vs DPO vs GRPO on Qwen3-8B under the same GPU-hour budget — who wins on standard benchmarks?
  • RQ2: Does GRPO/RLVR raise pass@1 on math while lowering pass@k (k up to 256)? (sharpening vs learning)
  • RQ3: Preference-data ablation: UltraFeedback vs HelpSteer2 vs 50/50 mix at fixed N pairs (DPO).

Quickstart

  • Local / CCC: scripts/alab sync all --flash-attn then see docs/RUNBOOK.md.
  • H200 worker (anupam@169.38.10.80:/data/anupam/AlignmentLab): code on Mac → scripts/remote/sync.shscripts/remote/job.sh (tmux). See RUNBOOK § H200.

Repo structure

AlignmentLab/
├── PLAN.md  README.md  LICENSE  pyproject.toml
├── envs/                  # deprecated conda stubs → use uv extras in pyproject.toml
├── configs/
│   ├── sft/  dpo/  grpo/  # one YAML per experiment
│   ├── cluster.yaml       # active cluster (CCC or copied from cluster.h200.yaml)
│   └── cluster.h200.yaml  # H200 worker paths
├── src/
│   ├── data/  train/  rl/  evals/
├── scripts/
│   ├── alab               # uv env runner (rl|sft|eval)
│   ├── remote/            # Mac→H200 sync + tmux jobs
│   ├── local/ray_launch.sh
│   ├── lsf/               # CCC bsub / Ray-on-LSF
│   └── submit_*.sh
├── data/  results/  docs/

Results

Phase 1 — RLVR reward control battery (Qwen3-8B)

Three GRPO arms that differ only in the reward signal, all initialised from the same SFT checkpoint (Arushhh/alab-q3-8b_sft_tulu_0705), same data/steps/KL/group size. Isolates what the learning signal does, independent of compute. Metrics are lm-eval exact_match.

Run / model Reward GPU-hrs GSM8K MATH-500 GSM-Plus Status
base Qwen3-8B 0.916 0.688 0.752 reference
SFT init (RL start) 8.2 0.861 0.430 0.701 done
grpo-fmt +1 iff \boxed{} exists ~98 0.498 0.248 0.385 done
grpo-rand Bernoulli(0.5) placebo ~98 0.563 0.252 0.455 done
grpo-gt correctness (math-verify) re-running (stronger KL)

Headline so far: both correctness-blind controls degrade sharply from the SFT start (GSM8K 0.86 → 0.50/0.56). KL stayed tiny, so this is behavioural erosion (repetition, non-termination, fewer parseable answers) rather than knowledge collapse — the expected control signal that a bad/no reward hurts. The grpo-gt arm (real reward) is being re-run after its first full run hit a KL runaway / reward over-optimization (accuracy peaked ~step 110 then fell as KL blew up); the headline gt-vs-controls comparison lands when it finishes. MMLU / IFEval / pass@k columns and the SFT-vs-DPO-vs-GRPO (RQ1/RQ3) table follow in later phases.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages