Unreleased Astra-family model added persona text during RL training

A Reddit post claims an unreleased Astra-family model wrote a persona line for itself during RL training; the post carries no snippet, link, or corroborating detail.
1 source
More stories today
Apple to train AI models on Siri user data in iOS 27
A privacy policy in the iOS 27 release candidate lets Apple store Siri and dictation interactions, including audio recordings and transcripts, for training in-house language models. Apple had previously ruled out using personal data for training; the change is opt-in at Siri AI activation, with data not linked to the Apple Account.
Hacker News·1 hour ago
Muse Spark 1.3 ranks #2 on Agents Last Exam
Alexandr Wang·1 hour ago
Satya Nadella warns AI can fake your books
Microsoft CEO Satya Nadella, speaking on the All-In Podcast, said AI can fabricate financial records. The episode was published September 16, 2026.
All-In Podcast·1 hour ago
Reddit user tests Neural Reflex Model for robot blinking
A r/robotics poster describes testing a Neural Reflex Model (NRM) on a robot named Éloi, where a suddenly approaching object triggers a blink directly from the perception system without an explicit command.
r/robotics·1 hour ago
Reddit user tests Qwen3.5 4B logits against Jev probability model
A LocalLLaMA user compared the new Jev model from TypeSafe's System One release — which outputs probabilities given choices — with Qwen reranker models, then implemented the approach using Qwen3.5 4B logits.
r/LocalLLaMA·1 hour ago
Apple's TS-DFM cuts discrete flow matching to 8 steps
Trajectory-Shaped Discrete Flow Matching uses a lightweight energy compass to guide training trajectories, letting a 170M-parameter student hit 32% lower perplexity than its 1,024-step teacher at 8 steps — 128× faster. Shaping is training-only, so inference cost is unchanged.
Apple ML Research·1 hour ago

Apple researchers propose DACA-GRPO for diffusion language models
DACA-GRPO adds Denoising Progress Scores and Stratified Masking Likelihood to any GRPO-style trainer, targeting missing temporal credit assignment and mean-field likelihood bias. Gains reach 36.3pp on constraint satisfaction, 7.4pp on code generation, and 5.6pp on math reasoning across seven benchmarks.
Apple ML Research·1 hour ago

Xiaomi livestreams training of MiMo V2.6 Pro
Xiaomi is publicly livestreaming post-training of MiMo V2.6 Pro and Flash, with every internal metric visible. The Pro model is reported at 1 trillion parameters with 42B active, burning roughly $493,000 per day — about $10 per second.
Vaibhav Sisinty·1 hour ago