Skip to content

Guide: Optimizing Cache Hit Rates

Cache hits are served instantly from memory, while compute requires processing. Maximizing your cache hit rate improves response time.

Note: With per-request pricing, cache hits and misses cost the same to you. However, higher cache rates improve your experience with faster responses.

The cache performs its own normalization, but you can improve hit rates further by pre-processing inputs on your client:

  • Trim whitespace: " hello " → "hello"
  • Lowercase: "Hello" → "hello" (if case doesn’t matter for your tools)
  • Remove punctuation: "hello!" → "hello"
const cleanQuery = rawInput.trim().toLowerCase().replace(/[^\w\s]/gi, '');

Keep tool descriptions stable across requests.

  • Don’t dynamically generate descriptions with timestamps or random IDs.
  • Sort tool parameters consistently.
  • If your tools don’t change between requests, use Standard Cache for the best hit rates.
ModeBest For
StandardStatic tool sets, high-volume apps, cost-sensitive workloads
AdvancedDiverse phrasing, multilingual input, dynamic tool sets

Standard mode gives the highest hit rates when your tools don’t change. Advanced mode catches more matches when users phrase the same intent in many different ways.

See Advanced Cache for details.

If you anticipate common queries (e.g., “Start”, “Menu”, “Help”), you can warm the cache by making scripted calls to the API during deployment or maintenance windows. This ensures your most frequent queries are served from cache from the start.

The cache automatically learns canonical forms of queries over time. The more traffic your app handles, the higher your hit rate becomes — without any action on your part.

You can accelerate this by using /v1/correct to teach the system the right answers for common queries. Corrections are stored in your Memory Banks and the corrected results get cached automatically.

When using Classification Sets, the classify endpoints unlock semantic caching — similar inputs return cached results. To maximize classify cache hits:

  • Use Classification Sets instead of inline classes. Inline classes only get exact-match caching.
  • Normalize your input text consistently before sending.
  • Stick to one quality tier per use case — different tiers produce separate cache entries.

When using inline classes, only exact-match caching is available. Fuzzy matching is too risky for classification accuracy when the class set isn’t stable.

Cache is automatically cleared when you modify a Toolset or Classification Set. To avoid unnecessary cache rebuilds:

  • Don’t update Toolsets or Classification Sets unless the content actually changes.
  • Renaming without changing tools/classes does not clear the cache.
  • Batch your changes — make all edits at once rather than one tool at a time.

When the cache identifies the correct tool but can’t extract a parameter value from the query, it normally falls back to compute. You can avoid this by setting "default" values on your tool parameters in the JSON Schema.

How it works:

  1. Cache hit identifies the tool.
  2. Regex extraction runs on the query for each parameter.
  3. For any parameter where extraction fails:
    • If the parameter has a "default" value → use it (no compute needed).
    • If the parameter is required with no default → fall back to compute.
    • If the parameter is optional with no default → omit it.

Example schema:

{
"type": "object",
"properties": {
"amount": { "type": "number" },
"currency": { "type": "string", "default": "USD" }
},
"required": ["amount", "currency"]
}

If the query is “send 50 dollars” and regex extracts amount: 50 but can’t determine currency, the system uses the default "USD" instead of requiring compute.

Observability: When defaults are used, the response metadata includes defaults_applied — an array of parameter names that were filled from defaults. The source remains "cache".

Notes:

  • Default values are not validated against schema constraints (type, enum, etc.) — ensure your defaults are valid.
  • Set defaults on parameters that have a sensible fallback (e.g., currency, locale, units) rather than parameters that are always user-specific.