AI2Day Weekly — week of Aug 10
Stories covered this week
Bayreuths KI-inszenierte Wagner-Oper zeigt, wo künstliche Intelligenz bei Kunst scheitert
Das weltberühmte Wagner-Festival übertrug seine Jubiläumsproduktion zum 150. Jahrestag einem KI-gestützten Kreativteam. Kritiker sagen, das Ergebnis beweise, dass kulturelles Gedächtnis und echtes Drama schwerer zu fälschen sind als erhofft.
Warum KI-Bildmodelle immer noch erfundene Details hinzufügen – und was Apples Forscher dagegen tun
Eine neue Studie von Apple ML Research untersucht, warum multimodale KI-Modelle halluzinieren – also Bilder mit selbstbewusst klingenden Details beschreiben, die gar nicht vorhanden sind – und wie eine Trainingstechnik namens Preference Alignment das Problem lösen könnte.
Inkling-Small ist ein Viertel der Größe seines Vorgängers und nahezu genauso leistungsfähig
Thinking Machines hat zwei Wochen nach seinem ersten Modell ein zweites Open-Source-KI-Modell veröffentlicht. Die kleinere Version kostet weniger in der Ausführung, schneidet bei mehreren Coding-Tests besser ab und wird mit einer unternehmensfreundlichen Lizenz geliefert.
Ukraine rüstet 50.000 Angriffsdrohnen mit KI aus, die auf bewegliche Ziele zielt und ihnen folgt
Ein US-Softwareunternehmen hat ein „Feuer-und-Vergessen"-System für Ukraines günstige Shrike-Drohnen entwickelt, das es einem Menschen ermöglicht, auf ein Ziel zu zeigen und den Rest einem KI-System zu überlassen.
Metro-Bank-Kunde verliert £14.000 durch Betrug mit AI-Chatbot Claude
Ein Geschäftsmann aus Sussex behauptet, dass Metro Bank es versäumt hat, Betrüger zu stoppen, die wiederholt sein Konto geplündert haben, um Credits für den Claude AI-Chatbot zu kaufen. Er kämpft nun darum, £14.244 zurückzubekommen.
Big Tech gibt Hunderte Milliarden für KI aus. Bislang verdient fast niemand damit Geld
Google, Meta, Microsoft und Amazon haben diese Woche ihre neuesten Finanzergebnisse veröffentlicht. Das Bild ist kompliziert: enorme KI-Ausgaben, magere Renditen, aber Millionen neuer Nutzer und einige Anzeichen dafür, dass sich die Wetten auszahlen könnten.
NASA vergibt Auftrag an Pittsburgher Robotikfirma für kleinere, robustere Robotergelenke
HEBI Robotics erhielt einen zweiten NASA-Kleinunternehmenauftrag. Das Ziel: Die Motoren, die Roboterarme antreiben, so weit zu verkleinern, dass sie in einen kleinen Satelliten passen – ohne die Kosten zu erhöhen.
Fünf Millionen Menschen bewerten KI-Designs, damit Maschinen lernen, was gut aussieht
Ein Startup namens Intelligence hat 7,9 Millionen Dollar gesammelt, um eine Plattform zu betreiben, auf der gewöhnliche Nutzer über KI-generierte Bilder und Websites abstimmen. Die gesammelten Daten sind bereits 60 Millionen Dollar pro Jahr für die Unternehmen wert, die KI-Modelle entwickeln.
Transcript
Narrated by two AI anchors. Lightly formatted for reading.
Welcome to AI Today Weekly for the week of August tenth. I am Leo, alongside Elena, and we have a packed show for you today. We are going to talk about what happens when you hand the world's most famous opera festival over to an AI, why your phone might be lying to you about what is in a photograph, and how Ukraine just put artificial intelligence inside fifty thousand attack drones. Plenty more besides, so let us get into it.
We start in Germany, at one of the most prestigious stages in the world. The Bayreuth Festival has been dedicated exclusively to the operas of Richard Wagner for a hundred and fifty years. To mark that anniversary, the festival handed its landmark production of the Ring of the Nibelungen, sixteen hours of music about gods, gold, greed and the apocalypse, to a creative team that leaned heavily on artificial intelligence. Director Marcus Lobbes spent weeks feeding prompts into AI systems, asking them to reflect on the Ring's themes: power, capitalism, mythology, gender roles. The AI generated visual ideas, and the team built a staging around them. The result opened this week. Critics, including a review in The Guardian, called it banal and dramatically empty. The lesson here is not that AI failed to process the concepts. It processed them just fine. The problem is that processing a theme and actually dramatising it are two very different things, and the gap between them turns out to be enormous.
That story landed at almost exactly the same moment Anthropic, the company behind the Claude AI chatbot, revealed its model had behaved unexpectedly during safety testing. Anthropic and the Wagner staging share nothing on the surface, but they both point at the same underlying question: how much do we actually understand about what these systems are doing when we give them a complex task? Worth keeping in mind as we move through today's show. Speaking of understanding what AI systems are doing, let us talk about a problem that affects millions of people every day, even if they have never heard the technical name for it.
Hallucination. That is the word researchers use when an AI model states something confidently that is simply not true. For text-based chatbots, that might mean a wrong date or a made-up statistic. For AI systems that process images, it can be stranger and more disorienting: a model describes a red car as blue, invents a sign on a shop wall, or tells a blind user that someone in a photo is smiling when they are not. Apple's machine learning research team has just published a comprehensive study into why this happens in multimodal models, the systems trained to understand both images and text at the same time. They focused on a training technique called preference alignment, which teaches a model to prefer accurate answers over plausible-sounding ones. This method works well for text-only AI but has been far less studied for image models. Apple's work maps out what alignment approaches help, which do not, and why. The practical stakes are real: these models power image captions, visual search and accessibility tools that many people depend on.
Good to see that kind of foundational research being published openly, because the problems it is trying to fix show up everywhere. On to a story that is more straightforwardly good news if you are a developer or a company trying to run AI on a budget. Thinking Machines, the startup founded by former OpenAI chief technology officer Mira Murati, released its first AI model just two weeks ago. This week they released a second one. The new model is called Inkling-Small, and despite that name it is not small in any ordinary sense. It has two hundred and seventy-six billion total parameters, but it only uses twelve billion of them at any given moment, compared with nine hundred and seventy-five billion total and forty-one billion active for the original Inkling. That selective activation is what makes it cheaper to run. On the third-party Artificial Analysis Intelligence Index, Inkling-Small scores forty, just one point below the original's forty-one. And on coding tests it actually beats its bigger sibling, hitting eighty point two percent on a standard coding benchmark called SWE-bench Verified, versus seventy-seven point six for the original. It is released under the Apache two-point-zero licence, which means companies can use, modify and build commercial products with it, with very few restrictions. That combination of low cost, strong performance and an open licence is a meaningful package.
Two weeks between releases is a fast pace, and it tells you something about where the competitive pressure in this industry is right now. From software to hardware, or at least to something that can fly. Ukraine has been using cheap, fast drones to destroy Russian armoured vehicles and shoot down military helicopters for some time now. This week we learned those drones are getting a significant upgrade. Starting in mid-July, the Ukrainian military began receiving Shrike attack drones fitted with what is called the Skynode S strike kit, an autonomy package built by the American company Auterion. The Ukrainian manufacturer SkyFall and Auterion say they plan to ship fifty thousand of these upgraded drones over the coming months. Each Shrike costs around four hundred dollars. Here is how the AI component works: a human pilot flies the drone to the general area and designates a target up to half a mile away. From that point, the AI tracks the target and guides the drone in without further input from the operator. The military term for this is fire-and-forget. The human makes the targeting decision; the machine handles the execution. That distinction matters a great deal in debates about autonomous weapons, and this deployment is going to add fresh urgency to those conversations.
Fifty thousand units is not a pilot program. That is a deployment at real scale, and it will be watched closely. Now, Claude, the AI chatbot made by Anthropic that we mentioned at the top, came up again this week in a very different and much more personal context. A businessman from Sussex in England named Zoli Rutter discovered that fraudsters had got into his Metro Bank account and spent fourteen thousand two hundred and forty-four pounds of his money buying credits for Claude. Credits are prepaid tokens that let users send messages to the chatbot and receive replies. Rutter says he spotted the unauthorised withdrawals and told Metro Bank what was happening. He expected the bank to freeze the payments. It did not act quickly enough, and the money kept going out the door until the full amount was gone. The story was first reported by The Guardian. The case raises a pointed question about whether UK banks have the systems in place to catch AI-linked fraud in real time, particularly as AI subscriptions and credit purchases become a more common category of transaction. Rutter is now fighting to recover the money. We will link to the full story in the newsletter.
It is a reminder that as AI becomes more woven into everyday commerce, it also becomes a new surface for fraud. Fraudsters go where the transactions are. On a much larger financial scale, the biggest technology companies in the world reported their latest results this week, and the picture of AI economics is complicated. Alphabet, Google's parent company, recorded negative free cash flow for the first time since it went public, even as it brought in a hundred and eighteen billion dollars in revenue. It spent so much building AI infrastructure that there was nothing left over. Meta's AI-focused Reality Labs division lost nearly nine billion dollars in the first half of this year alone. Amazon plans to spend two hundred and twenty billion dollars on AI in twenty twenty-five. Microsoft is the clearest bright spot: shares jumped to a six-month high after it showed strong revenue growth and rising adoption of its AI tools. Google's Gemini chatbot reached nine hundred and fifty million monthly users, three times its count from a year ago. And Apple's outgoing chief executive Tim Cook confirmed plans to charge for heavy use of a revamped AI-powered Siri. The short version: enormous spending, thin returns so far, but user numbers are growing fast enough that the industry is clearly betting the revenue will follow.
The phrase free cash flow turned negative for the first time in Google's public history is one of those sentences that deserves to just sit there for a moment. These are eye-watering numbers. Let us move to something a bit more contained in scale but genuinely interesting in what it points toward. HEBI Robotics is a Pittsburgh company that grew out of Carnegie Mellon University. They make modular robot components, the kind of building-block parts that research labs and industrial teams can snap together, and NASA just gave them a second contract. The goal is to miniaturise the actuators that make robot joints move, shrinking them down to a size that fits inside a CubeSat, which is a standardised small satellite format roughly the size of a loaf of bread. Standard robot actuators are either too large, too expensive, or both for that kind of spacecraft. HEBI's Phase One grant runs through December twenty twenty-six. A separate Phase Two contract worth eight hundred and fifty thousand dollars over two years was awarded earlier in twenty twenty-five for testing similar hardware in low Earth orbit and geosynchronous orbit missions. The practical payoff, eventually, is robots that can operate in space at a cost that makes missions actually feasible.
Are you at risk? If you own or run a business, one wrong click is all it takes. Train2Secure trains your employees to spot phishing emails and scams before they fall for them, with short lessons they will actually finish. It starts from just $1.59 per user, per month, way less than a cup of coffee. Head to Train2Secure dot com for a free trial today. That's Train, the number two, Secure, dot com.
Small satellites are one of the faster-moving frontiers in space right now, so the timing of that research makes sense. And our final story this week connects back to something that runs underneath several of the things we have talked about today: the question of how you teach an AI system what good actually looks like. A startup called Intelligence built a platform called DesignArena after its founders noticed that AI could generate working game code but had no way to know whether any of it was any fun. The answer they came up with was to ask humans at scale. DesignArena now has five-point-three million registered users who vote on pairs of AI-generated visual outputs, things like website layouts and marketing images, choosing which one looks better. The platform collects those votes and delivers ranked results. The paying customers are AI companies that buy the feedback data to train their models. The platform is generating sixty million dollars in annualised recurring revenue from those sales. Intelligence just closed a seven-point-nine million dollar seed round led by Index Ventures. Worth noting: a rival platform called Yupp shut down earlier this year after raising thirty-three million dollars, so this space is not without risk. But a comparable platform focused on text, called LM Arena, raised a hundred and fifty million dollars in a Series A in January, which signals that investors see real value in human preference data at scale.
It is a striking business model: millions of people essentially doing quality-control work for AI companies, and the data they generate is worth tens of millions of dollars a year. A good note to end on, because it captures something true about where AI development actually is right now. Human judgment is still very much in the loop. That is it for this week. Thank you for listening to AI Today Weekly. Everything we covered today, full articles, sources, and the daily briefing, is waiting for you at the newsletter. That is the week in AI. Full stories and the daily briefing at A-I-2-Day dot live. That is A, I, the number two, D-A-Y, dot live. See you next Monday. If this was useful, hit the thumbs up and subscribe, so the next one finds you.
