Pinterest cuts AI costs 90% by optimizing vision model layer
TL;DR. Pinterest reduced AI inference costs by 90% after streamlining its visual search model through a novel architectural change. - The company replaced a complex transformer layer with a simpler attention pooling mechanism in its vision transformer. - This optimization maintained model performance while significantly lowering computational demands and costs. - The new architecture enabled a tenfold increase in query capacity compared to the previous model.
- Pinterest optimized its visual search AI by gutting a frontier model's vision layer.
- The primary change involved replacing a transformer in the vision layer with a simple attention module.
- This architectural simplification resulted in a 90% cost reduction for AI inference.
- The improved model achieved a threefold increase in inference speed and 10x higher query capacity.
- Despite the drastic cost reduction, the AI model retained its core performance capabilities.
Sources
- Pinterest cut AI costs 90% by gutting a frontier model's vision layer — venturebeat.com