Is the API You're Paying For Real?

The most common relay scam is paying premium-model rates while the backend quietly hands you a cheaper model or an older version — commonly called "model substitution" or "downgrading." Below, we cover how these tricks actually work, followed by a checklist you can run yourself with zero tools, before you ever commit money.

Short on time? Just remember these two

  • A price that doesn't add up, with no explanation for where it comes from, is usually a problem — this is the single strongest signal in this space.
  • ② To really confirm, testing it yourself with your own key is the most direct approach. A single test is noisy — run it a few times and look at the trend.

5 forms of "model substitution / downgrading"

1

Cheap-model substitution

You're billed at Claude or GPT-5 rates, but the backend is actually running something far cheaper like GPT-4o-mini, Qwen, or DeepSeek. Hard to catch in casual chat, but it shows up fast on complex reasoning or multi-step tasks.

2

Silent version downgrades

You requested the latest Sonnet 4.6 or Opus 5, but you're actually getting an older, cheaper version. Benchmark performance won't match, even though the model name looks identical.

3

Trimmed context / capped output

Advertised as supporting long context, but overly long inputs get silently truncated — or max output length gets quietly capped to save cost. Shows up as the model "forgetting" earlier content, or responses cutting off mid-sentence.

4

Spoofed response fields

Wraps a third-party model's output to look exactly like the official response structure — faking the model name, stop_reason, and billing fields. Nearly impossible to catch with the naked eye or a quick script.

5

Intermittent swaps

Only switches to the cheap backend during peak hours or specific windows, and behaves normally the rest of the time. So "it checked out when I tested it" doesn't mean it's clean all the time — this is the hardest variant to catch.

The free manual self-check checklist

No tools required — run these 5 steps before you top up.

1

Compare the price and see if they can explain where it comes from

Compare a specific model's per-unit price against the going market rate. Note: cheap by itself isn't a problem — pricing from legitimate channels like subscription slicing, regional discounts, or IDE promotional quotas is often genuinely low. The real warning sign is "priced well below everyone else, and can't explain why." This step has the best payoff for the least effort.

2

Run fixed test questions side by side

Prepare a handful of moderately difficult questions you know well, run them through a provider you trust and the one you're testing, and compare style, depth, and reasoning quality — are they in the same tier?

3

Long-context needle test

Feed in a sufficiently long block of text, then ask a question at the end that's only answerable if it actually processed the earlier content. If it "forgets" or the content got cut off, the context may have been silently trimmed.

4

Version-capability needle test

Ask it something the claimed version should handle well, but that an older or cheaper model would struggle with, and check whether its actual ability matches the version it claims to be.

5

Retest at different times

Don't test just once. Test again at a different time of day and over a different connection — this is specifically how you catch the "normal most of the time, swapped during peak hours" pattern.

What manual testing can't catch, and a data point worth knowing

Manual self-checks can only tell you what "looks" right — they can't see billing-layer fingerprints or output statistical signatures, the harder-to-fake evidence underneath, so they're inherently subjective and can be gamed. Model substitution isn't a rare unlucky occurrence, either — a 2026 empirical paper from the CISPA Helmholtz Center for Information Security ("Real Money, Fake Models," arXiv:2603.01919) found that nearly half (45.83%) of 24 tested relay endpoints failed model identity verification. It's worth taking seriously before you commit a large deposit.

An honest caveat: any single test run has noise, and "this looks off" is not proof of fraud — repeat testing across different times gives you a far more reliable read on the trend. HowToken's one-click online verification tool is currently in development.

Channel types explainedPitfalls to avoidCompare prices across providers