Frontier LabsLiquid AI's New Draft Models Make Its LFMs Up to 3.2x Faster, No Quality Trade-off
A technique called speculative decoding lets a small helper model do the heavy lifting so the main model just checks the work. The result: dramatically faster output on everything from a data-centre GPU to a MacBook.