16.6points
How far benchmark scores moved when the only change was which OpenRouter provider served the model.
LessWrongAuthentication for open-source models
Some providers serve a cheaper model, a compressed copy or a broken setup under the name you picked. authenticated.si checks each call against the real model and tells you what you got.
deepseek/deepseek-v4-pro
via example-host
Authenticated
96% confidence · 1 credit
Hover the certificate to see it under UV.
You pick DeepSeek V4 Pro and get a smaller, cheaper model billed under the same name.
The right model squeezed to 4-bit. Easy prompts look fine; long code and hard math start to slip.
The right model with its context window cut from 128k to 32k tokens, or tool calls that come back as broken JSON.
16.6points
How far benchmark scores moved when the only change was which OpenRouter provider served the model.
LessWrongMay 52026
A strict client caught a private-inference gateway serving Qwen3.5-122B in place of DeepSeek-V3.1.
awesome-private-inferenceA one-time test isn't enough. Developers testing Kimi providers found some pass at light load and slip when their GPUs are busy. Hacker News thread
Your terminal or agent sends a short set of test prompts to the endpoint you already use.
We compare the answers with reference answers we made by running the real model ourselves.
Six signals are combined into one verdict with a confidence level, and the check costs a credit.
A sneaker legit check looks at the stitching, the size tag and the sole, because no single detail settles it. We weigh six signals for the same reason.
Every model family splits text into tokens differently, so the bill for the same prompt gives a different model away.
With randomness turned off, the real model words the same answer the same way every time.
ReferenceThe capital of Australia is Canberra.
EndpointCanberra is the capital of Australia.
A compressed copy starts out matching, then wanders from the reference partway through a long answer.
Tricky math and code prompts are where 4-bit copies lose points first.
We hide a code word at the start of a very long prompt. A provider that quietly cut the context window can't find it.
We ask for a function call. The real setup returns clean JSON; a broken one returns text that your code can't parse.
Valid: {"name":"get_weather","arguments":{"city":"Delhi"}}Broken: get_weather(city=DelhiEvery signal matches the reference for the model you picked.
Same model family, but answers drift partway through and hard questions slip.
Token counts and fixed answers point to a different model than the one on your bill.
The model matches, but the context window is cut short or tool calls break.
$ authenticated check deepseek/deepseek-v4-pro \
--endpoint $PROVIDER_URL --depth quick
✓ token count 612 / 612
✓ fixed answers 10 / 10
✓ drift none in 200 tokens
✓ hidden word found at 128k
AUTHENTICATED 96% confidence 1 credit
# before each batch of work
POST https://api.authenticated.si/v1/checks
{
"model": "deepseek/deepseek-v4-pro",
"endpoint": "$PROVIDER_URL",
"depth": "quick"
}
# 200 OK
{ "verdict": "authenticated", "confidence": 0.96 }
Pay per check. Credits come in prepaid packs.
Three depths. Quick, Standard and Deep.
Any OpenAI-compatible endpoint. Including routers like OpenRouter.
We plan to serve open models ourselves. Our own endpoints will run the same checks in public, so you can hold us to the standard we hold everyone else to.
Kartikeya Sharma co-founded Dype, a sneaker authentication service, in 2019. authenticated.si brings the same legit check to AI models, built by Bhag Labs, Inc.
Checks use our own test prompts, sent to the endpoint you name. Your own conversations don't pass through authenticated.si.
No. It's a weighted verdict with a confidence level, built from six independent signals. Like a legit check, it's strong evidence, and we tell you how strong.
No. We test the endpoint from the outside, the same way you use it, so it works with any provider that offers an OpenAI-compatible API.
We start with a short list of popular open models whose reference answers we generate ourselves, and add more as each reference run finishes.
You buy a pack of credits and each check spends some, depending on depth. There's no subscription. Prices go live at launch, and people on the early access list hear first.