I Cut Duplicate Payments for LLM Answers, Revealing Limits of Embeddings for Prompt Deduplication
agents embeddings
| Source: Mastodon | Original article
A developer discovered that duplicate LLM responses can incur double charges, revealing that embeddings alone are insufficient for deduplicating prompts.
A developer working on an open‑source Rust‑based LLM agent discovered that the system was paying for identical answers twice – once when an automated retry fired and again when a user double‑clicked the same request, causing two workers to query the model in parallel. The root cause was a naïve reliance on embedding‑based deduplication: identical vectors were not enough to recognise re‑phrased prompts that still required the same response.
The fix came in the form of a “semantic cache” that stores full completions and matches new prompts against them using a similarity threshold. As the accompanying technical note explains, setting the threshold too low returns incorrect answers, while a too‑high setting reduces the cache to an exact‑match store, negating its purpose. The developer also added hit‑rate instrumentation from day one, turning the cache from a guess‑work add‑on into a measurable cost‑saving layer.
Why it matters: In production LLM deployments, even modest duplicate traffic can inflate operating costs and latency. Traditional caching, which only matches exact prompt strings, fails when users ask the same question in different wording – a common pattern in support bots and internal tools. By moving beyond embeddings to semantic similarity of full prompts, operators can cut token spend and improve response times without sacrificing answer quality.
What to watch next: The community is now debating best practices for threshold tuning and hit‑rate monitoring, as highlighted in recent posts on semantic caching. A complementary “embedding cache” – a low‑risk optimisation that avoids re‑embedding unchanged text – is being rolled out in parallel, per a July 2026 guide. As we reported on 24 February 2026 in “PromptCache Part I: Stop Paying Twice for the Same LLM Answer”, the industry is rapidly converging on layered caching strategies. Expect further tooling, open‑source libraries, and benchmark data to emerge over the coming months, helping developers balance cost, latency, and answer fidelity.
Sources
Back to AIPULSEN