In this interview with Impact Newswire, Shruti Koparkar, NVIDIA’s AI Inference Product Lead, explains how NVIDIA’s Blackwell GPUs and Dynamo inference software are helping Pinterest scale AI-powered visual discovery by processing more images with lower latency. She discusses the engineering challenges of multimodal AI, including the computational demands of image processing, memory bandwidth, caching and routing, and how precomputed PinCLIP visual embeddings reduce repeated computation. Koparkar says Pinterest’s benchmarks show a 25-fold increase in visual context per request, an 85-fold improvement in response startup and a 7.3-fold reduction in end-to-end latency.

Imagine asking an artificial intelligence assistant to help redesign your living room. You show it a photograph of your sofa, a collection of lamps you have saved, several hundred pictures of rooms you like and a handful of rugs you are considering. You want it to understand not just what is in each image, but which pieces belong together, which colours complement one another and what might work in your own home.
For a human, this is an exercise in visual judgment. For an AI system, it is a computational problem that begins before the first word of an answer appears.
Each image may need to be loaded, decoded, processed and translated into a representation the model can interpret. Multiply that work across hundreds of images, then repeat it as a conversation develops, and the machinery behind a seemingly simple shopping assistant begins to face a difficult trade-off: how much visual information can it consider without making users wait?
Pinterest, the visual discovery platform, is attempting to push that boundary with NVIDIA, whose Blackwell processors and Dynamo inference software now underpin a new layer of the company’s multimodal AI infrastructure. Pinterest says the system allows its AI-powered Pinterest Assistant to process 25 times more visual context per request. In benchmark testing, the company reported an average improvement of about 85 times in response startup and 7.3 times in end-to-end latency when using precomputed visual representations instead of repeatedly processing raw images.
The figures point to a broader shift in the economics of artificial intelligence. As companies move beyond text-based chatbots toward systems that interpret images, video and other forms of data, the challenge is no longer simply to build a more capable model. It is to deliver that capability quickly and economically enough to support millions of users.
Pinterest says its platform serves 640 million monthly active users and handles more than 80 billion searches each month. Its visual discovery business gives the company a particular incentive to make AI understand images at scale, potentially connecting inspiration more directly to product discovery and shopping.
The industry has been moving in the same direction. Stanford University’s 2025 AI Index found that the cost of querying a model performing at approximately GPT-3.5’s level on a standard language benchmark fell from $20 per million tokens in November 2022 to $0.07 by October 2024, using one of the less expensive models in its comparison. That decline illustrates how much inference costs can change through advances in models, hardware and deployment techniques. It does not, however, establish that every AI workload has become cheaper at the same rate.
For NVIDIA, the Pinterest deployment also illustrates why faster processors alone may not be enough. Blackwell supplies the computing capacity; NVIDIA Dynamo coordinates how requests move through the serving system; the open-source vLLM engine executes the model; and Pinterest’s own visual embeddings reduce the need to process images repeatedly.
Those contributions matter because the headline improvements do not all measure the same thing. NVIDIA says preliminary tests found that Blackwell B200 delivered more than twice the latency improvement over its Hopper-generation hardware for Pinterest Assistant. The much larger startup-time gain, by contrast, is associated with a change in how Pinterest represents and serves visual information, alongside the software and infrastructure used to process it. The results are company-reported benchmarks, not an independently verified, hardware-only comparison.
Shruti Koparkar, NVIDIA’s AI Inference Product Lead, spoke with Faustine Ngila about the bottlenecks in visual AI, the role of inference orchestration and the engineering choices behind Pinterest’s reported gains. The questions examine where the improvements originate and what other companies should consider before expecting comparable results.
Here is the interview excerpt:
1. NVIDIA says Blackwell was particularly well suited to Pinterest’s vision-language workloads because they are heavily weighted toward prefill. What did NVIDIA identify as the specific bottlenecks in Pinterest’s existing infrastructure, and which Blackwell capabilities were most important in addressing them?
Pinterest’s engineering team characterized these bottlenecks in their own production workloads and the two teams co-engineered to address them. Vision-language inference is more demanding than text-only workloads. Before the model generates a single token, every image has to be downloaded, preprocessed, encoded, and converted into tokens. Large, variable visual contexts and growing KV-cache from multi-turn conversations add to the pressure, creating real contention between prefill and token generation.
Blackwell B200 is built for exactly this. Higher low-precision compute throughput, greater memory bandwidth, and a second-generation Transformer Engine accelerate prefill heavy workloads. In Pinterest’s preliminary Pinterest Assistant benchmarking, B200 cut latency by more than 2x compared to Hopper.
2. Dynamo is positioned as a major part of NVIDIA’s inference stack rather than simply another optimization tool. What does Dynamo do differently at the systems level, particularly in routing, disaggregating inference workloads and managing GPU resources, that enabled Pinterest to achieve these gains?
Dynamo is a distributed inference orchestration and serving layer, not a single optimization.
Three capabilities matter most here. Disaggregated serving splits image encoding, prefill and decode into separately scalable pools, so Pinterest can tune each phase against a latency target instead of forcing one set of GPUs to do all three. KV-aware routing sends each request to the most optimal worker with the most relevant cached context and available capacity, rather than simply the first free worker, which avoids recomputation. KV cache offloading tiers that context across GPU, CPU memory and NVMe so it’s not evicted and re-computed in long multi-turn conversations.
Dynamo is also open source, Kubernetes-native and inference-engine agnostic allowing Pinterest to run vLLM, its inference-engine of choice underneath it.
3. NVIDIA reports an 85x improvement in response startup and a 7.3x reduction in latency for consumer-facing AI at Pinterest. How did NVIDIA isolate the contribution of the Blackwell hardware from software optimizations in Dynamo, vLLM and Pinterest’s own serving architecture? What should customers understand about where those gains actually come from?
Those benchmarks measured changes to Pinterest’s serving approach on a platform built with Blackwell, Dynamo, and vLLM.
The key breakthrough was enabling Pinterest to reuse visual information it had already computed, substantially reducing the work required for each request. Delivering that efficiently required the hardware, inference software, and Pinterest’s architecture to work together. Customers should view these results as evidence of what optimizing the full application on an accelerated platform can achieve.
4. Pinterest says a single request can involve thousands of images. What is actually happening inside the inference pipeline when that request arrives? How does Dynamo decide what gets encoded, cached, routed and processed, and where are the biggest computational bottlenecks?
A request can contain text alongside images. Raw images must be loaded, decoded, and encoded into representations the model can understand. When Pinterest already has embeddings for that content, the pipeline reuses them, avoiding repeated image processing.
Dynamo coordinates the serving pipeline, allowing image encoding, prompt processing, and response generation to scale independently. Its multimodal KV-aware routing helps direct requests to workers that have cached processing of shared visual content, reducing recomputation. vLLM executes the model and manages the attention state used during inference.
The biggest bottlenecks are the computation required to process large visual inputs before the first response token and the memory bandwidth needed to manage that context. Multi-turn conversations increase that pressure. NVIDIA Blackwell’s low precision compute and high memory bandwidth capabilities and Dynamo’s multimodal orchestration provide a scalable foundation for serving these rich visual experiences efficiently.
5. Pinterest claims Dynamo allows Pinterest Assistant to process 25 times more visual context per request. What specifically enabled that 25x increase? Is the gain primarily coming from KV-cache management, projection embeddings, batching, routing, parallelism, or a combination of these techniques?
The specific 25x result is primarily tied to Pinterest’s use of precomputed PinCLIP projection embeddings. In its benchmark, a request containing 250 images represented as PinCLIP visual embeddings achieved latency comparable to a pixel-based request containing roughly 10 images. By avoiding repeated processing of raw pixels, the system can devote substantially more of its latency and compute budget to reasoning over visual context.
Dynamo’s routing, disaggregated serving and cache-management capabilities help make that larger context practical and scalable in production, while vLLM executes the model and Pinterest’s custom projector converts the embeddings into the model’s visual-token space.
6. Pinterest’s internal image encoder reportedly delivers up to a 44x improvement in end-to-end latency by serving precomputed visual embeddings rather than repeatedly processing raw pixels. What are the engineering trade-offs here?
Precomputed embeddings reduce latency by avoiding repeated image processing. The trade-off is added engineering to support those embeddings across the serving stack,
7. Pinterest reports more than 2x lower latency on B200 compared with Hopper for Pinterest Assistant and an overall 85x faster response startup. How should we distinguish the hardware contribution from the software contribution? In other words, how much of this improvement comes from Blackwell itself versus Dynamo, vLLM, routing and Pinterest’s model optimizations?
The more-than-2× result compares B200 with Hopper and represents preliminary latency gains on Blackwell for Pinterest Assistant, while the roughly 85× improvement in average response startup reflects the move from raw-pixel processing to precomputed PinCLIP embeddings within Pinterest’s serving architecture, Dynamo, and vLLM.
Stay ahead of the Stories shaping our world. Subscribe to Impact Newswire and join our
WhatsApp Channel for updates on global tech, business, and innovation—all in one place.
Dive deeper into the future with the Cause Effect 4.0 Podcast, where we explore the ideas, trends, and technologies driving the global AI conversation.
Got a story to share? Contact Us to reach a global audience with Impact Newswire.
Faustine Ngila is the AI Editor at Impact Newswire, based in Nairobi, Kenya. He is an award-winning journalist specializing in artificial intelligence, blockchain, and emerging technologies.
He previously worked as a global technology reporter at Quartz in New York and Digital Frontier in London, where he covered innovation, startups, and the global digital economy.
With years of experience reporting on cutting-edge technologies, Faustine focuses on AI developments, industry trends, and the impact of technology on society.
Discover more from Impact Newswire
Subscribe to get the latest posts sent to your email.



