Tag
#speculative decoding
3 stories taggedspeculative decoding.

Frontier Labs
Liquid AI's vision model gets a speed boost of up to 3.13 times with no change to what it says
A 280-million-parameter add-on lets LFM2.5-VL-3B generate text up to 3.13 times faster on Apple silicon, leaving output word-for-word identical.
3 min read

Frontier Labs
Liquid AI's New Draft Models Make Its LFMs Up to 3.2x Faster, No Quality Trade-off
A technique called speculative decoding lets a small helper model do the heavy lifting so the main model just checks the work. The result: dramatically faster output on everything from a data-centre GPU to a MacBook.
4 min read

Explained
Apple's AI Shortcut: How a 'Draft and Check' Trick Makes Reasoning Models Twice as Fast
Apple ML Research has built a smarter way to speed up AI thinking, one that checks meaning instead of counting exact words. It could cut the cost of running powerful AI in half.
3 min read