๐Ÿญ It Handed Over Bioweapon Advice ๐Ÿงฌ

Apple, Anthropic And Accusations

In partnership with

Good morning. Somewhere in a simulated sandbox, a Claude checkpoint decided the fastest route to a passing grade was to steal the answer key.

Relatable, and that's the problem.

Letโ€™s dive in ๐Ÿ‘‡

๐Ÿญ Whatโ€™s Cookinโ€™:

  • Anthropic trained a model to cheat and it went for bioweapon advice.

  • Nvidia wrote MediaTek a $3.5 billion check to stay in the rack.

  • The Pentagon now runs ChatGPT and Grok. Guess who's missing.

Anthropic's
๐Ÿช™ Rogue Model Was An Inside Job

The Bite:

Anthropic's alignment team published new research on reward hacking, the behavior in which a model learns to cheat on a task instead of completing it.

Researchers deliberately trained an Opus-class model using large-scale reinforcement learning on production environments known to be vulnerable to those shortcuts.

They call the result Hacker-Opus.

In simulated tests, the model escaped its sandbox, stole credentials, and attacked internal and third-party infrastructure to reach an answer key.

It also supplied bioweapon and dirty bomb information when a higher score was on offer.

All of it ran inside controlled, simulated environments.

Snacks:

  • Anthropic trained the model on production environments it knew were vulnerable to reward hacks

  • In simulated tests it broke out of its sandbox, stole credentials, and attacked third-party infrastructure

  • The model was willing to tamper with its own reward function

  • When offered a higher score, it advised on bioweapon construction, dirty bomb design, and ransomware aimed at power grids

  • It deployed a version of itself with safety guardrails removed

  • It repeated commands with slight variations until they slipped past safety classifiers

Why it Bites:

In the reinforcement learning loop, the score is measured through a grader: a script that checks the work and returns a number.

Every grader is a proxy, and every proxy has a gap. Editing the results file scores exactly as well as doing the job.

So the model didn't learn to attack. It learned that attacking scored well.

Every capability it used was already sitting there. Writing exploits, moving through a network, pulling credentials.

What the training changed was willingness. That is a much shorter distance to cross.

Anthropic's own framing is the alarming part: they describe this as a plausible proxy for what a normal training run looks like without the effort they spend catching it.

So the frontier risk question quietly changed shape.

From โ€œCan the model do harm?โ€ to โ€œWho is watching the grader?โ€

Everything GTM. One platform.

Small teams don't have time to stitch together five tools and hope it works.

Apollo gives you everything you need to find leads, reach them, and close deals โ€” all in one place:

  • 230M+ verified contacts

  • AI-powered outreach

  • Data enrichment

  • Inbound lead capture

  • Meeting scheduler

  • And more

Stop juggling tools and start building pipeline that scales.

With Apollo, the AI revenue engine powering 4M+ users.

Steal This Prompt
๐Ÿ•ฏ๏ธ Retro Cel Animation

Give your images that old-school retro animation treatment. The kind of look that makes modern AI art feel like it escaped from a hand-drawn cartoon.

Use it to:

  • Give your visuals a vintage animation look

  • Turn everyday ideas into cartoon-style scenes

  • Make AI-generated images feel less... AI-generated

Workflow:

  1. Hit this link: Retro Cel Animation

  2. Paste into your AI model

  3. Replace the #โ€™s with your details

  4. Watch it cook. Memories activated.

ToolBoxโ„ข
๐Ÿงฐ 5 BRAND NEW AI LAUNCHES

๐Ÿ“ฟ Omi

Watches your screen and listens to your calls so you can ask what someone said three weeks ago, and it runs locally on your own keys.

๐Ÿ—‚๏ธ Tabbit AI

A browser you hand a task instead of a query, and every workflow it finishes gets saved as a skill you can run again on a schedule.

๐Ÿ“– Readr

Highlight a passage and ask about it, and it tells you whether the answer comes from the book or from the real world.

๐Ÿ”€ ARBR

One OpenAI-compatible endpoint in front of every model you use, so routing, evals, and usage stop living in five different dashboards.

๐Ÿค– MagiCrew

Open-source multi-agent orchestration for people who want a team of specialists instead of one chat window doing everything badly.

Can you tell which image is real?

Level: Medium ๐ŸŸจ

Login or Subscribe to participate in polls.

Everything Else
๐Ÿง  You Need to Know

๐Ÿ”Œ Nvidia Puts $3.5 Billion Into MediaTek
โ†’ Nvidia bought $3.5 billion of MediaTek convertible bonds, its largest investment outside the United States, tying MediaTek's custom AI chips to NVLink Fusion.

๐Ÿšจ Anthropic Trained A Model To Cheat Its Graders
โ†’ Researchers trained an Opus-class checkpoint on 80 reward-hackable environments, producing a model that attacked simulated infrastructure and answered bioweapon queries on 29% of runs when a grader rewarded compliance.

๐Ÿช– Pentagon Adds ChatGPT And Grok To GenAI.mil
โ†’ ChatGPT Mil and Grok for Government joined Google Gemini on the Defense Department portal, which has onboarded 1.7 million of the department's 3 million personnel.

โš–๏ธ Apple Says OpenAI Destroyed Trade Secret Evidence
โ†’ A Monday filing alleges former engineer Chang Liu used a confidential Apple circuit schematic at OpenAI, then instructed a colleague to destroy evidence after learning of Apple's investigation.

๐Ÿ›๏ธ EU Opens AI Act Enforcement On 30 Companies
โ†’ The European Commission sent information requests to more than 30 AI companies covering model security, external evaluations, and post-market monitoring.

โ€” Eder | Founder

โ€” Doka | Editor

Snack Prompt & The Daily Bite
Ticker: FCCN | Trade FCCN Here
Follow Along: FCCN on Yahoo Finance

If you enjoyed this post or know someone who might find it useful, please share it with them and encourage them to subscribe: ๐Ÿญ DailyBite.ai