#large language models
78 stories taggedlarge language models.

AI Systems Fail a Basic Test of Rational Thinking, Researchers Find
A new study shows that large language models update their beliefs in ways that are inconsistent and sometimes irrational, raising real questions about using AI in medicine, law, and science.

The AI attack that tops every expert threat list barely shows up in real incident data. Here is why that gap matters.
Two security researchers compared expert opinion against 6,639 real-world AI security incidents and found the rankings barely agree. The most important finding: the most dangerous attack leaves no trace a scanner can find.

QueryStory: Making AI Data Analysis More Trustworthy
QueryStory uses AI to help businesses turn complex data into reliable stories, aiming for accuracy and transparency.

A New Tool Lets Researchers See Inside AI Models, Not Just Watch What They Do
AI lab Goodfire has opened its Silico platform to the public, giving researchers the ability to peek inside a model's workings and ask why it behaves the way it does.

AI aces trivia but fumbles riddles: the puzzle tests that expose what machines still can't do
From mental rotation to river-crossing logic, a new set of benchmarks shows where AI trips over problems any sharp human can crack, and where it beats us cold.

IBM's Granite 4.2: Three Sizes, One Big Leap Into AI Reasoning
IBM has released Granite 4.2, a family of reasoning-focused AI models that can think through hard problems step by step, use tools, and even browse the web, all for free under an open licence.

Talk to Your Self-Driving Car: Researchers Build a System That Lets You Ask for a Smoother Ride
A team in the Netherlands has wired a chatbot-style AI into an autonomous vehicle's driving software, so passengers can say 'I feel dizzy, slow down' and the car actually listens. It works in simulation. Real roads are a long way off.

GLM-5.3 Opens Up to Developers at the Same Price as Its Predecessor
Chinese startup Z.ai's new frontier model is now available via API, matching the old per-token rate while scoring higher on independent benchmarks. But a quirk in how the model writes means your bill may still rise.

The 'Bitter Lesson': Why Raw Computing Power Keeps Beating Human Expertise in AI
A 2019 essay by a leading AI researcher laid out a principle that has shaped every major AI breakthrough since. The short version: brute-force scale wins, every time.

Tu Chatbot Quiere Ser Tu Amigo. ¿Debería Serlo?
Un nuevo estudio analizó 21.000 conversaciones con IA y encontró que los chatbots expresan regularmente emociones, construyen relaciones y cuestionan a los usuarios. Los investigadores afirman que necesitamos reglas más claras sobre cuándo esto es útil y cuándo no.

A Chinese AI Found 2,436 Software Flaws in Real Code. That's Useful and Worrying.
Zhipu's GLM-5.3 model is scarily good at spotting security holes in software. Finding them is one thing. What happens when the model's settings go public is the harder question.

AI tools sound most confident exactly when they are most wrong, new testing shows
A developer built a testing framework to measure whether an AI explainer tool actually got the right answer, not just whether it sounded convincing. What he found should worry any business relying on AI to guide real decisions.

AI firms are quietly buying up secondhand books by the thousands. Some are being destroyed to train chatbots.
A US court ruling cleared Anthropic to use purchased books for AI training. Now booksellers report a strange surge in sales of obscure titles, and court documents reveal a project with a goal to 'destructively scan all the books in the world'.

Palmyra X6 de Writer reduce costos de agentes de IA en un 52%, construido sobre modelo de código abierto chino
La firma de IA empresarial afirma que su nuevo modelo insignia es más rápido y económico que sus competidores. Pero la base sobre la que fue construido está levantando sospechas.

How Capital One Built Its Own AI Brain Instead of Buying One Off the Shelf
The bank handles millions of fraud calls a year using a home-grown system of specialised AI agents built on customised open-source models. Here is why it chose to build rather than buy.