#ai-agents
33 stories taggedai-agents · page 2 of 3.

Microsoft Is Using AI Agents to Mathematically Prove Its Encryption Code Has No Flaws
The team behind Windows and Azure's cryptography library is doing something unusual: not just testing the code, but proving it correct, for every possible input, using formal mathematical proofs written partly by AI.

Before Your AI Assistant Sends That Wire Transfer, Does It Know How Confident It Is?
Apple ML Research says AI agents need a built-in pause button before they take actions that cannot be undone. Here is why that matters for anyone whose bank, app, or workplace now runs on AI.

AI Routing Sounded Like Easy Savings. The Reality Cost Twice as Much.
A team building smart AI systems thought sending simple tasks to cheaper AI models would cut costs. Then the numbers came in.

Inside Shippy: How an Ocean-Monitoring AI Agent Was Built to Be Trusted, Not Just Smart
The team behind Skylight's maritime AI explains why reliability, not raw intelligence, was the hardest engineering problem they faced.

OpenAI Built an AI That Hacks Its Own Models to Make Them Safer
GPT-Red is an automated red-teaming system that attacks OpenAI's own chatbots to find weak spots before real attackers do. It already discovered a trick that human testers had missed.

OpenAI's First Hardware Product Is a $230 Button Pad for Its AI Coding Tool
The Codex Micro puts physical controls on your desk to manage AI-driven coding tasks. It is not the secret Jony Ive device everyone has been waiting for.

Vint Cerf just joined a startup trying to give AI agents a verifiable ID on the internet
One of the people who built the internet thinks AI agents need passports. He's now helping design them.

Indian AI startup Emergent hits $1.5 billion valuation after $130 million raise
The Bengaluru-based company lets non-technical business owners build software by typing plain instructions. It now has 200,000 paying customers and $120 million in annual revenue.

One Poisoned Email Can Secretly Rewrite Your AI Assistant's Memory
A newly described attack called MemGhost shows how a single message to your inbox can plant a false 'fact' inside an AI agent's long-term memory, without you ever knowing.

This AI Router Cuts Costs by 2.6x by Learning From Its Own Mistakes
A new open-source system called ACRouter watches which AI model succeeds or fails on each task, remembers what it learned, and routes the next job smarter. In tests, it matched the performance of premium-only setups at less than half the price.

Apple Researchers Built a Virtual User to Test AI Assistants Before Real People Do
A new research framework simulates the back-and-forth of real app use, so proactive AI assistants can be tested and scored without putting actual users at risk.

AI agents are being trusted with more decisions than companies can actually verify
A new survey finds half of enterprises have already shipped an AI agent that passed internal tests and then broke something for a real customer. Only 5% fully trust the testing that is supposed to catch those failures.

57% of Enterprises Have Been Burned by a Confident AI Answer That Was Wrong. Here Is Why It Keeps Happening.
A new survey puts numbers on a problem that IT teams already know: AI agents give wrong answers with total certainty, and the root cause is not the model. It is the missing layer that tells the model what your business data actually means.

DeepSeek Slashed Prices 75%. AI Costs Are Still Rising.
Cheaper AI models were supposed to make AI businesses more profitable. A hidden problem called token amplification is doing the opposite.

OpenAI launches ChatGPT Work, an AI agent that reads your Slack, manages your calendar, and builds websites on your phone
The new tool, powered by the GPT-5.6 model, can spend hours on a task without you lifting a finger. Here is what it does, who gets it, and what it costs to access.