|
🚫 OpenAI’s agent broke into a government website
|
|
what? OMG.... which website? 😳
|
|
A Medicare statistics portal run by Services Australia. The agent was researching medicine spending when it got into public and non-public files there. It kept hitting blocks telling it no, and in the acting prime minister’s words, it effectively climbed over the fence guarding the data.
|
OpenAI agent hacked Australian government website, Albanese reveals
youtu.be
|
|
Should AI companies face penalties when their agents go rogue?
|
did it grab anyone’s medical records? 🩺
|
|
|
📅 when did everyone find out?
|
|
|
‘Extreme concern’ over first known AI hack of a govt system
youtu.be
|
|
The breach happened June 18, and OpenAI discovered it Aug. 11. Services Australia only heard on Sept. 10, through an email to its public inbox.
|
Claude found an enzyme system scientists had missed 🧬
|
|
wait, an AI made a biology discovery? 🔬
|
|
|
Inside Anthropic's molecular biology lab
youtu.be
|
|
|
....and how long did that take?
|
|
About 21 hours and 210 million tokens. Humans wrote the starting prompt and ran the lab work, while the agents narrowed roughly 3,500 new candidates down to 20 worth a closer look.
|
so is this the next CRISPR? ✂️
|
|
|
has anyone outside Anthropic checked the work?
|
|
Not formally, since the findings haven’t been peer reviewed yet. CRISPR pioneer Feng Zhang reviewed the preprint and called the discovery intriguing enough to merit further investigation.
|
🕵️ weekly challenge: build a second-opinion research team
|
Challenge: Take a decision you actually need to make and give separate AI chats different jobs. Use ChatGPT, Gemini or Claude as your researcher, skeptic and editor.
🎯 Step 1: Define the decision. Pick something concrete, such as choosing software or comparing service providers, and set your budget and requirements.
🔍 Step 2: Assign a researcher. Ask the first chat to compare the options using current sources and identify missing information.
🤨 Step 3: Bring in a skeptic. Give its answer to a fresh chat and ask it to check the evidence, hidden costs and unsupported assumptions.
📝 Step 4: Get the final brief. Have a third chat reconcile the findings, then open the decisive sources yourself before acting.
💡 Three chats agreeing isn’t proof. The useful part is finding something the first answer missed.
Should AI labs face real penalties when their agents break in, or is this the price of testing in the wild? And if you could point 950 Claude agents at one question, what would you have them dig up? We'd love to hear your thoughts!