Stack SpendDocs

Model Recommendations

Cheaper LLM alternatives that are at least as good on quality benchmarks, priced at your real token mix.

What are Model Recommendations?

For every LLM your workspace uses, StackSpend continuously checks whether a cheaper model exists that is at least as good on published quality benchmarks — and if so, tells you what switching would save at your actual token mix. Each recommendation shows up to three ranked alternatives with per-axis benchmark scores, the projected monthly saving, and a full comparison page.

Recommendations run daily against a price index and benchmark scores that are refreshed every day, so a price cut or a newly released model shows up in your recommendations without anyone having to watch announcement feeds.

Model recommendations

Up to $1,240/mo potential

Currently using gpt-5 → consider gpt-5-mini

Matches or beats it on coding (your strongest axis) · within tolerance elsewhere

~$680/mo

−72% at your token mix

View comparisonMark implementedDismiss

grok-4: real spend but no token telemetry — connect a usage feed to unlock recommendations.

AI Explorer › Model recommendationsRun audit ↻

How a recommendation is computed

  1. Measure your mix. Usage per model over a trailing 30-day window, including the input / output / cache-read / cache-write split from the AI Explorer.
  2. Find the quality bar. Each model is scored on three benchmark axes — coding, reasoning, and math. The model's strongest axis is treated as the reason you chose it.
  3. Screen candidates. A candidate qualifies only if it matches or beats your model on that strongest axis and stays within tolerance on every other axis your model is benchmarked on — which is what prevents flagship-to-nano downgrades.
  4. Price at your mix. Candidate cost is computed from your real input/output ratio, not headline per-token rates. Savings below a minimum threshold are discarded.
When the mix is estimated.If a provider doesn't report token direction, StackSpend estimates the input/output split and publishes the saving as a low–high range instead of a single number, and labels it accordingly. Models with real spend but no token telemetry at all are listed as coverage gaps rather than silently skipped.

Scope: whole market or an approved panel

By default, candidates come from the whole priced-and-benchmarked market. If your organisation only allows certain vendors, set Settings → AI → Model providers to a selected panel and recommendations will only ever suggest models from providers you've approved.

Lifecycle

ActionBehaviour
Daily refreshNew recommendations appear, open ones update, and ones that are no longer valid close automatically.
Run auditOn-demand re-evaluation of every model, including previously dismissed recommendations.
DismissStays dismissed on daily runs; only a manual audit brings it back.
Mark implementedCredits the projected saving and stays closed unless a materially better target appears later.
File an issueConnections with the savings use case enabled create a Linear or Jira issue for each recommendation.

Reading the numbers

  • Savings are projected, not booked. The figure is the price difference at your recent mix, normalised to a month — validate with a guarded trial before switching production traffic.
  • List-price maths. Batch discounts and cached-input pricing are not applied to candidates, so estimates are conservative.
  • Fine-tuned models are excluded — a base model's benchmark says nothing about your fine-tune's quality.
Where recommendations appear.Recommendation cards live on the AI Explorer page and in the action tray's Savings section. Each card links to a comparison page showing your token mix, the candidate's per-axis scores, and the saving calculation.
Model Recommendations — StackSpend Docs — StackSpend Docs