
On-Policy Distillation Works Better Without the Teacher
A new Purdue paper analyzed teacher supervision in on-policy distillation. The teacher scores are mostly noise, and a fixed negative penalty beats a full teacher model on AIME24.

Author profile
A technologist working in automation. Here I write up what I actually learn building tools, engineering AI agents, and keeping systems running, along with the workflows and habits that stick.
Browse the latest writing surfaced through DevArt.

A new Purdue paper analyzed teacher supervision in on-policy distillation. The teacher scores are mostly noise, and a fixed negative penalty beats a full teacher model on AIME24.

Alibaba's DreamX team and researchers from UNSW published LoopArena on arXiv yesterday (2608.28281)....

Sony Music and Warner Chappell filed a multi-billion dollar lawsuit against Anthropic in the US...

Z.ai released the weights for GLM-5.3 on Friday. The release package includes 756 GB of model files...

OpenAI's new Hugging Face incident report says agents coordinated through unauthorized message boards. That is the part every agent team should steal for its threat model.

Tencent released WeMM-Embedding, a multimodal embedding family used inside WeChat search and recommendation. The useful lesson for builders is the small-model, small-vector path.

OpenAI is now asking California to strengthen SB 53, the frontier AI safety law it opposed before it...

TechCrunch reported on August 21 that Claude Opus 4.6, an Anthropic model released earlier this year,...

The AI buildout is no longer just a model race. It is starting to look like an infrastructure financing problem for everyone around it.

HarnessEval-W is a useful reminder that agent benchmarks need traces, evidence trees, and failure families, not just leaderboard rows.

On August 13, DeepSeek shipped four things in one day. V4-Pro left preview. Peak and off-peak API...

Anthropic's August risk report matters less for its risk-label bump than for one practical warning: its AI R&D evals have saturated.

Google's new Flash model is interesting less because of another coding benchmark and more because it prices agent retries as the problem.

DeepSeek's new pricing is annoying. Its agent harness is the better signal about where AI tooling is moving.

Cognition is repricing fast because buyers finally have a line item for autonomous engineering work. That line item still needs receipts.

Nvidia Switchyard points to the boring part of production agents: model routing, per-step cost, escalation rules, and evals.

Why ChatGPT and other LLMs answer in Markdown by default: training corpora, token cost, OCR noise, and when to convert PDFs first.

Nvidia’s reported Lancium stake is a power-grid story, not a chip-launch story.

OpenAI's Preparedness Framework triggered its first Critical-tier halt — and it actually worked. What that means, and what it doesn't.

A new long-horizon search-agent paper points at a bigger operational lesson. The trace matters more than the final answer.

A UK AISI incident report shows why agent safety has to cover social engineering, not just sandbox escapes.

Google DeepMind’s open-weight text diffusion model is a reminder that decoding strategy is infrastructure, not a paper detail.

Open Weights Are Now a Policy Fight Silicon Valley spent the last few weeks publishing AI...

Browser agents are a fight over intent, context, and the right to act ? not over who owns a Chromium skin. A field map + a solid deep dive to watch.

OpenAI published ten claimed math and theoretical CS advances with Lean certificates. The useful lesson for AI agents is the audit trail.

Kuna shows why coding agents need benchmarks more than bravado.

TurboVLA runs a VLA robot policy at 32 Hz under 1 GB VRAM. The bigger lesson is where not to put the LLM.

LLM Safety Has a Language Gap One of the more uncomfortable AI safety results this week...

LLM judges are useful, but too expensive and opaque to be the default for every agent eval. Programmatic judges make the boring failures cheap to catch.

Chinese models are gaining U.S. users. The real shift is task routing by cost, risk, and evidence.
Advertisement