AI Research

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Medium Severity Global
Date Occurred Aug 19, 2026 17:54 UTC
Event Type AI Research
Source arXiv
Recorded Aug 20, 2026
Full Description

arXiv: Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial success. We diagnose this mismatch on fixed responses fro

AI Intelligence Layer

Mentioned Models

Qwen

AI Categories

ethics performance
Event Metadata
  • ID #25074
  • Type AI Research
  • Region Global
  • Severity Medium
  • Indexed Aug 20, 2026