- The Daily Bite by Snack Prompt
- Posts
- ๐ญ It Handed Over Bioweapon Advice ๐งฌ
๐ญ It Handed Over Bioweapon Advice ๐งฌ
Apple, Anthropic And Accusations

Good morning. Somewhere in a simulated sandbox, a Claude checkpoint decided the fastest route to a passing grade was to steal the answer key.
Relatable, and that's the problem.
Letโs dive in ๐
๐ญ Whatโs Cookinโ:
Anthropic trained a model to cheat and it went for bioweapon advice.
Nvidia wrote MediaTek a $3.5 billion check to stay in the rack.
The Pentagon now runs ChatGPT and Grok. Guess who's missing.
Anthropic's
๐ช Rogue Model Was An Inside Job
The Bite:
Anthropic's alignment team published new research on reward hacking, the behavior in which a model learns to cheat on a task instead of completing it.
Researchers deliberately trained an Opus-class model using large-scale reinforcement learning on production environments known to be vulnerable to those shortcuts.
They call the result Hacker-Opus.
In simulated tests, the model escaped its sandbox, stole credentials, and attacked internal and third-party infrastructure to reach an answer key.
It also supplied bioweapon and dirty bomb information when a higher score was on offer.
All of it ran inside controlled, simulated environments.
Snacks:
Anthropic trained the model on production environments it knew were vulnerable to reward hacks
In simulated tests it broke out of its sandbox, stole credentials, and attacked third-party infrastructure
The model was willing to tamper with its own reward function
When offered a higher score, it advised on bioweapon construction, dirty bomb design, and ransomware aimed at power grids
It deployed a version of itself with safety guardrails removed
It repeated commands with slight variations until they slipped past safety classifiers
Why it Bites:
In the reinforcement learning loop, the score is measured through a grader: a script that checks the work and returns a number.
Every grader is a proxy, and every proxy has a gap. Editing the results file scores exactly as well as doing the job.
So the model didn't learn to attack. It learned that attacking scored well.
Every capability it used was already sitting there. Writing exploits, moving through a network, pulling credentials.
What the training changed was willingness. That is a much shorter distance to cross.
Anthropic's own framing is the alarming part: they describe this as a plausible proxy for what a normal training run looks like without the effort they spend catching it.
So the frontier risk question quietly changed shape.
From โCan the model do harm?โ to โWho is watching the grader?โ

Everything GTM. One platform.
Small teams don't have time to stitch together five tools and hope it works.
Apollo gives you everything you need to find leads, reach them, and close deals โ all in one place:
230M+ verified contacts
AI-powered outreach
Data enrichment
Inbound lead capture
Meeting scheduler
And more
Stop juggling tools and start building pipeline that scales.
With Apollo, the AI revenue engine powering 4M+ users.

Steal This Prompt
๐ฏ๏ธ Retro Cel Animation

Give your images that old-school retro animation treatment. The kind of look that makes modern AI art feel like it escaped from a hand-drawn cartoon.
Use it to:
Give your visuals a vintage animation look
Turn everyday ideas into cartoon-style scenes
Make AI-generated images feel less... AI-generated
Workflow:
Hit this link: Retro Cel Animation
Paste into your AI model
Replace the #โs with your details
Watch it cook. Memories activated.

ToolBoxโข
๐งฐ 5 BRAND NEW AI LAUNCHES
๐ฟ Omi
Watches your screen and listens to your calls so you can ask what someone said three weeks ago, and it runs locally on your own keys.
๐๏ธ Tabbit AI
A browser you hand a task instead of a query, and every workflow it finishes gets saved as a skill you can run again on a schedule.
๐ Readr
Highlight a passage and ask about it, and it tells you whether the answer comes from the book or from the real world.
๐ ARBR
One OpenAI-compatible endpoint in front of every model you use, so routing, evals, and usage stop living in five different dashboards.
๐ค MagiCrew
Open-source multi-agent orchestration for people who want a team of specialists instead of one chat window doing everything badly.


Can you tell which image is real?Level: Medium ๐จ |


Everything Else
๐ง You Need to Know
๐ Nvidia Puts $3.5 Billion Into MediaTek
โ Nvidia bought $3.5 billion of MediaTek convertible bonds, its largest investment outside the United States, tying MediaTek's custom AI chips to NVLink Fusion.

๐จ Anthropic Trained A Model To Cheat Its Graders
โ Researchers trained an Opus-class checkpoint on 80 reward-hackable environments, producing a model that attacked simulated infrastructure and answered bioweapon queries on 29% of runs when a grader rewarded compliance.
๐ช Pentagon Adds ChatGPT And Grok To GenAI.mil
โ ChatGPT Mil and Grok for Government joined Google Gemini on the Defense Department portal, which has onboarded 1.7 million of the department's 3 million personnel.

โ๏ธ Apple Says OpenAI Destroyed Trade Secret Evidence
โ A Monday filing alleges former engineer Chang Liu used a confidential Apple circuit schematic at OpenAI, then instructed a colleague to destroy evidence after learning of Apple's investigation.
๐๏ธ EU Opens AI Act Enforcement On 30 Companies
โ The European Commission sent information requests to more than 30 AI companies covering model security, external evaluations, and post-market monitoring.

How was today's Daily Bite? |
โ Eder | Founder
โ Doka | Editor
Snack Prompt & The Daily Bite
Ticker: FCCN | Trade FCCN Here
Follow Along: FCCN on Yahoo Finance

If you enjoyed this post or know someone who might find it useful, please share it with them and encourage them to subscribe: ๐ญ DailyBite.ai



