CrofAI Fraud Exposed: 'World's Cheapest AI Inference' Turns Out to Be an OpenRouter Wrapper Scam

CrofAI's 'cheapest inference' was an OpenRouter wrapper with up to 20x markups — then it vanished.
CrofAI, which claimed to be the world's cheapest AI inference provider, was exposed by independent researcher Kendell as a pure OpenRouter wrapper that silently rerouted requests to cheap models like GLM, Qwen, and Kimi — with markups up to 20x. Its advertised 'proprietary greg model family' was fabricated, as the founder admitted in private messages. Five 'fix' attempts only obscured OpenRouter's fingerprints without building any real infrastructure. Hardware claims were mathematically impossible based on public specs. Faced with wire fraud allegations, CrofAI denied, then backtracked, staged a fake 'team takeover,' and finally deleted all online presence.
A company called CrofAI, which billed itself as the "world's cheapest AI inference provider," has been unmasked by an independent investigator. Its supposedly proprietary inference engine never existed — the service was merely an OpenRouter wrapper that silently rerouted user requests to smaller, cheaper models while marking up prices by as much as 20x. Confronted with the allegations, CrofAI first denied everything, then backtracked, and three hours later wiped all traces of its online presence. This is a cautionary tale about chasing dirt-cheap tokens.

Prices Too Good to Be True
CrofAI (operating under the domains crof.ai and nahcrof.com) had one core selling point: access to all the latest models at prices often significantly lower than the cheapest options available on OpenRouter. Its founder claimed to run a custom inference engine that enabled these rock-bottom token prices, and even mocked competitors for charging more due to their "skill issues."
This narrative was enormously appealing to cost-conscious developers. But according to a detailed technical analysis published by researcher Kendell on their blog (kendell.dev), the reality was straightforward: "CrofAI is an OpenRouter wrapper that quietly routes your requested model to a cheaper or weaker model."
In other words, users believed they were calling top-tier models and paying accordingly — but what they actually received were cheap substitutes, with CrofAI pocketing the difference.
The Mechanics of a 20x Markup
The gap between advertised pricing and actual routing revealed by the investigation is staggering. Take the expensive kimi-k3 as an example: CrofAI sold it at $2 per million input tokens and $10 per million output tokens, but the requests were actually forwarded to GLM 5.3 Flash on OpenRouter. That translates to a 13.3x markup on input and a 20x markup on output.
Even more absurd was the so-called "proprietary" greg model family. The investigation found:
- greg-2-ultra actually routed to GLM 5.2
- greg-1-mini actually routed to Qwen 3.5 9B
- greg-2-super, greg-1, greg-1-super actually routed to Kimi K2.7 Code
All of these "proprietary models" were sold at significant premiums over the models actually being called. In private messages, CrofAI's founder admitted the greg family was entirely fabricated.
GLM 5.3 Flash, Qwen 3.5 9B, and Kimi K2.7 Code all belong to the open-source or low-cost commercial model tier, with inference costs far below flagship models. The GLM series is jointly developed by Tsinghua University and Zhipu AI; the Flash variant is optimized for high-throughput, low-latency scenarios and is typically priced at a fraction of a cent per million tokens. Qwen 3.5 9B is Alibaba's mid-size Qwen model — at 9 billion parameters, it can run on a single consumer-grade GPU with minimal marginal inference cost. Kimi K2.7 Code is a lightweight model from Moonshot AI trimmed specifically for coding tasks. These models differ fundamentally from top-tier flagships (like the full Kimi K3 or GPT-4-class models) in both capability and price, though similar naming conventions make confusion easy. OpenRouter, as an aggregation layer, publicly lists the actual billing price for each model — which is precisely how the investigator was able to calculate exact markup multiples by comparing CrofAI's charges against OpenRouter's transparent pricing.
Five "Fixes" That Only Hid the Fingerprints
Before publishing the exposé, the investigator gave CrofAI considerable time and advance warning, and documented five separate attempts by CrofAI to "fix" the problem of requests being proxied through OpenRouter. The results were telling: across all five attempts, the only thing CrofAI ever did was try to conceal OpenRouter's identifying fingerprints — while the requests continued to flow through OpenRouter regardless.
This behavioral pattern is itself compelling evidence. A company with a genuine proprietary inference engine would have no reason to go to such lengths to hide traces of a third-party service.
OpenRouter leaves identifiable "fingerprints" in HTTP response headers, error message formats, and specific field naming conventions — for instance, responses may carry x-openrouter-* headers, or platform identifiers embedded in specific fields of streaming output. The investigator confirmed the OpenRouter routing by capturing these fingerprints. CrofAI's five "fix" attempts included filtering response headers and modifying error return formats, but none of them severed the actual connection to OpenRouter. After each "fix," the investigator was still able to verify the routing path through deeper-level characteristics. This pattern of changing appearances while leaving the substance untouched became a central piece of evidence in the investigation report — a service provider running its own infrastructure simply would not care whether a third party's fingerprints were visible.
The Hardware Math Doesn't Add Up
Beyond the routing evidence, CrofAI's technical claims also fall apart at the hardware level.
CrofAI claimed to be running Kimi K3 on RTX Pro 6000 GPUs rented through Vast. But even using the extremely aggressive Q2_K quantization (what the industry jokingly calls a "lobotomy"-level compression), the model requires approximately 802 GiB of VRAM. The largest RTX PRO 6000 configuration available on Vast is 8 cards, totaling only 765 GiB — not enough to load the model at all.
Even more outlandish: the founder claimed he would run deepseek-v4-flash-0731 locally on a DGX Spark to "investigate" why his API was routing through OpenRouter. A single DGX Spark has only 128 GB of memory — also nowhere near sufficient to run that model. These self-contradictory technical details further cement the fictional nature of the entire narrative.
Q2_K quantization is one of the most aggressive compression schemes in the GGUF format, reducing model weights from the original 16-bit or 32-bit floating point to an average of approximately 2.5 bits. This can reduce VRAM requirements to roughly one-sixth of the original, but at the cost of significant model capability degradation — the gap in inference quality on sensitive tasks compared to full-precision models is substantial. The industry's "lobotomy" nickname is apt. Vast.ai is a GPU compute rental marketplace where users can rent consumer and professional graphics cards by the hour. The RTX PRO 6000 is NVIDIA's latest professional workstation card with 96 GiB of VRAM per unit. Even with 8 cards maxed out (765 GiB), the 802 GiB minimum requirement cannot be met — a gap provable from publicly available specs alone, with no black-box testing required.
From Denial to Disappearance
CrofAI's public relations response to the exposure was a farce from start to finish.
Initially, the founder announced the service was shutting down, acknowledged he could not deliver proprietary inference, and promised refunds to anyone who requested one. According to archived records, at approximately 4:30 AM UTC on September 15th, he published a blog post — later deleted — written in the voice of a "team," claiming that all of the founder's previous statements "were written under enormous pressure and described the situation worse than it actually was," and announcing that a new team had taken over and the service would resume within two weeks.
For good measure, CrofAI's Twitter account was also "taken over by the team," with every reply beginning with "Hey, Nathan here," claiming the founder had stepped back. This "fake team" performance lasted only a few hours. Perhaps the public simply wasn't buying another story from a serial liar, or perhaps repeated reminders from users that "wire fraud carries multiple charges" had an effect — either way, he ultimately deleted all online traces: nahcrof.com and crof.ai now return 404, the Twitter account has been deleted, and the /r/CrofAI subreddit has been set to private.
A Warning About Cheap Tokens
This incident is a wake-up call for the entire AI developer community. When a provider's prices are dramatically lower than every competitor in the market, and the explanation given is that everyone else lacks the technical skill — the rational response is suspicion, not celebration.
AI inference is fundamentally a capital- and compute-intensive business. Hardware and electricity costs set a hard floor on pricing. Any provider claiming to sustainably and substantially undercut that floor deserves technical verification — checking response fingerprints, comparing model behavioral characteristics, and auditing hardware feasibility. CrofAI was exposed precisely because of this kind of rigorous technical scrutiny.
For teams relying on third-party inference APIs, the question of "is the model you're paying for actually the model you think you're getting" is becoming an issue of trust that can no longer be ignored.
Related articles

Building Open-Source Video Editor Concat with Claude: A Free Alternative to CapCut
A developer used Claude to build Concat, an open-source CapCut alternative, in just three weeks. With nearly 10K downloads, it's a striking example of AI-assisted solo development.

The Overlooked Gems of Self-Hosting: Fun and Useful Services Nobody Talks About
From a Reddit thread, we explore overlooked self-hosted services — including video scraper RECLIP and data viz project WORLD MONITOR — and why fun services are so rare.

Bringing Software Quality Back to the Mainstream: A Reflection on Speed-First Culture
Why has software quality gone from default to luxury? Exploring technical debt, software bloat, and how to make quality the norm again in a speed-obsessed industry.