#multimodal AI
7 stories taggedmultimodal AI.

NeoMME Is a Tiny AI Model That Reads Documents Twice as Fast as Its Rivals
A new open-source model scans document pages as images, skips the usual heavy architecture, and still matches much larger systems on accuracy. Here is what that means if you work with PDFs.

Apple's AI Model Can Write Text and Draw Pictures at the Same Time, Using One Engine
A new research model from Apple blends language and image generation into a single, unified system. Here is why that matters for the next generation of AI assistants.

Meta's Muse Glimmer Is a 30-Billion-Parameter AI You Can Run on Your Own Machine
The new open-source model sees images, understands video, and can act as a personal assistant without sending your data to a cloud server.

Mistral's New Safety Tool Reads Your Rules and Applies Them Instantly
Shieldstral is a small, free AI model that screens text and images for harmful content using plain-English policies you write yourself. No specialist knowledge required.

Why AI Image Models Still Make Things Up, and What Apple's Researchers Are Doing About It
A new study from Apple ML Research digs into why multimodal AI models hallucinate, meaning they describe images with confident-sounding details that simply are not there, and how a training technique called preference alignment could fix it.

Inkling-Small Is a Quarter the Size of Its Predecessor and Nearly as Capable
Thinking Machines has released a second open-source AI model just two weeks after its first. The smaller version costs less to run, scores higher on several coding tests, and comes with a business-friendly licence.

Researchers Build a Test to See If AI Can Actually Summarise a 16-Minute Video
A new benchmark called LVSum reveals how badly today's AI video tools lose track of when things happen in long recordings.