Model Recommendations
Cheaper LLM alternatives that are at least as good on quality benchmarks, priced at your real token mix.
What are Model Recommendations?
For every LLM your workspace uses, StackSpend continuously checks whether a cheaper model exists that is at least as good on published quality benchmarks — and if so, tells you what switching would save at your actual token mix. Each recommendation shows up to three ranked alternatives with per-axis benchmark scores, the projected monthly saving, and a full comparison page.
Recommendations run daily against a price index and benchmark scores that are refreshed every day, so a price cut or a newly released model shows up in your recommendations without anyone having to watch announcement feeds.
Model recommendations
Up to $1,240/mo potentialCurrently using gpt-5 → consider gpt-5-mini
Matches or beats it on coding (your strongest axis) · within tolerance elsewhere
~$680/mo
−72% at your token mix
grok-4: real spend but no token telemetry — connect a usage feed to unlock recommendations.
How a recommendation is computed
- Measure your mix. Usage per model over a trailing 30-day window, including the input / output / cache-read / cache-write split from the AI Explorer.
- Find the quality bar. Each model is scored on three benchmark axes — coding, reasoning, and math. The model's strongest axis is treated as the reason you chose it.
- Screen candidates. A candidate qualifies only if it matches or beats your model on that strongest axis and stays within tolerance on every other axis your model is benchmarked on — which is what prevents flagship-to-nano downgrades.
- Price at your mix. Candidate cost is computed from your real input/output ratio, not headline per-token rates. Savings below a minimum threshold are discarded.
Scope: whole market or an approved panel
By default, candidates come from the whole priced-and-benchmarked market. If your organisation only allows certain vendors, set Settings → AI → Model providers to a selected panel and recommendations will only ever suggest models from providers you've approved.
Lifecycle
| Action | Behaviour |
|---|---|
| Daily refresh | New recommendations appear, open ones update, and ones that are no longer valid close automatically. |
| Run audit | On-demand re-evaluation of every model, including previously dismissed recommendations. |
| Dismiss | Stays dismissed on daily runs; only a manual audit brings it back. |
| Mark implemented | Credits the projected saving and stays closed unless a materially better target appears later. |
| File an issue | Connections with the savings use case enabled create a Linear or Jira issue for each recommendation. |
Reading the numbers
- Savings are projected, not booked. The figure is the price difference at your recent mix, normalised to a month — validate with a guarded trial before switching production traffic.
- List-price maths. Batch discounts and cached-input pricing are not applied to candidates, so estimates are conservative.
- Fine-tuned models are excluded — a base model's benchmark says nothing about your fine-tune's quality.