Running Qwen for 63 Hours to Tackle the Riemann Hypothesis: An Open-Source Model Autonomous Reasoning Experiment

Qwen 3.8 27B ran autonomously for 63 hours on a single RTX 3090 attempting to prove the Riemann Hypothesis.
A developer deployed Qwen 3.8 27B (4-bit quantized, 100K context) on a single RTX 3090 for 63 hours, generating over 50 million tokens in an attempt to prove the Riemann Hypothesis. While the model didn't crack this Millennium Prize Problem, the experiment revealed three notable properties: zero hallucinated answers throughout, continuous exploration of new strategies, and repeated self-correction. The full dataset is publicly available on Hugging Face. The experimenter argues the real value lies not in the result but in the methodology — a stress test of open-source model stability during long-duration autonomous reasoning, with plans to next try stronger models or multi-agent swarms on more tractable open problems.
A Bold Experiment: 63 Hours and 50 Million Tokens
A Reddit user shared an ambitious experiment: they ran Qwen 3.8 27B (4-bit quantized, with a 100K context window) autonomously on a single RTX 3090 GPU for 63 consecutive hours, generating over 50 million tokens — all in an attempt to prove one of mathematics' most famous unsolved problems: the Riemann Hypothesis.
The result was unsurprising: the model didn't solve this Millennium Prize Problem. But the experimenter argued that the value of this attempt wasn't in the final answer — it was in observing the model's internal working mechanisms, memory organization, and strategy evolution during prolonged autonomous operation.

The Riemann Hypothesis was proposed by German mathematician Bernhard Riemann in 1859. Its core claim is that all non-trivial zeros of the Riemann zeta function have a real part equal to 1/2. Deceptively simple in statement, it is deeply connected to the distribution of prime numbers and is a cornerstone of analytic number theory. The Clay Mathematics Institute lists it as one of the seven Millennium Prize Problems, offering a $1 million reward. To date, thousands of professional mathematicians have attempted to crack it — all without success. Its extreme difficulty lies in requiring an entirely new mathematical framework, not merely an extension of existing tools — which is precisely why brute-force AI reasoning falls short.
Qwen 3 is a series of open-source large language models released by Alibaba's Tongyi Lab, featuring hybrid reasoning modes (switchable between "thinking" and "non-thinking" modes) and reinforcement learning post-training capabilities. 4-bit quantization is a model compression technique that dramatically reduces VRAM usage by compressing model weights from 16/32-bit floats to 4-bit integers, enabling a 27B-parameter model to run on a single consumer-grade GPU (such as the RTX 3090's 24GB VRAM), at the cost of slight precision loss. The 100K context window means the model can "remember" approximately 100,000 tokens of history in a single pass — a key enabler of this extended autonomous reasoning session.
The Most Noteworthy Observations from the Experiment
The experimenter noted several surprising phenomena. First, over the entire 63-hour run, the model never fabricated a false answer — which is unusual for large language models that routinely hallucinate. Second, the model never stopped trying new approaches, continuously switching strategies and exploring different proof paths throughout its reasoning chain.
Even more interesting: the model proactively corrected its own mistakes multiple times. This self-correction capability is one of the most closely watched areas in current AI Agent research. It suggests that in an ultra-long context, autonomous loop environment, the model exhibited some degree of "metacognition" — the ability to look back at its own reasoning and identify flaws.
The experimenter published the complete process data on Hugging Face, including the model's internal "memory," generated code, adopted strategies, and other full records for the community to study and reproduce.
Metacognition in cognitive science refers to "thinking about thinking" — an individual's ability to monitor, evaluate, and adjust their own cognitive processes. In the LLM field, this concept is borrowed to describe a model's behavior of tracing back through its own reasoning chain, identifying logical contradictions, and actively correcting course. The autoregressive generation mechanism of traditional LLMs means they are fundamentally forward-predicting token by token, without a natural ability to "look back." Models like Qwen 3, which feature explicit Chain-of-Thought reasoning and reinforcement learning training, reward self-correction behavior during training — enabling the model to question and revise already-generated content during inference. This is one of the core breakthrough directions in current Agent reliability research.
Why Failing to Solve It Still Matters
To be clear-eyed about this: the Riemann Hypothesis is widely recognized as one of the hardest problems in mathematics, and is one of the Clay Mathematics Institute's seven Millennium Prize Problems. Problems of this nature require genuinely new mathematical insight — not brute-force computation or an ever-larger pile of tokens. The probability that a 27B open-source model solves it in 63 hours is essentially zero.
But the real value of this experiment lies in its methodological exploration: it demonstrates the stability boundaries of current open-source models during long-duration autonomous operation. Generating 50 million tokens continuously without crashing, falling into meaningless loops, or producing hallucinated answers is itself a stress test of model robustness.
For developers researching AI Agents, datasets like this are invaluable — they document how a model organizes memory, how it plans tasks without human intervention, and how it redirects itself when hitting a dead end. These are precisely the core questions that need to be understood in order to build reliable autonomous agents.
Next Steps: Agent Swarms and Stronger Models
The experimenter expressed considerable optimism about the possibility of "an open-source model-powered agent or agent swarm solving a Millennium Prize Problem within the next 12 months." While this prediction may be overly aggressive, the underlying technical direction is worth watching.
Their next plan is to use stronger open-source models (mentioning smarter models in the vein of the GLM series), or to assemble an agent swarm to collaboratively tackle a "simple but still open" mathematical or programming problem. This shift from a single Agent to multi-Agent collaboration aligns with the mainstream trend in current Agent research — improving overall reasoning quality through division of labor, debate, and cross-validation.
The experimenter also extended an open invitation to the community: if you have spare GPU resources, you could join in running multiple Agents as a swarm to collectively attempt a verifiable open problem.
Agent Swarms / Multi-Agent Systems refer to architectures where multiple AI agents work in parallel or in coordination, with each agent potentially handling different subtasks (e.g., literature retrieval, formula derivation, code verification), cross-checking results via shared memory or message passing. Compared to a single Agent, swarm architectures offer key advantages: parallel exploration of multiple proof paths, error filtering through "debate" mechanisms, and reduced context pressure on individual agents through division of labor. Frameworks like AutoGen, CrewAI, and LangGraph already provide toolchains for building such systems. However, multi-agent systems also introduce new complexity: communication overhead between agents, consistency maintenance, and error propagation are all active research challenges.
Implications for the Open-Source AI Community
While this experiment was a personal endeavor, it reflects a characteristic spirit of exploration in the open-source AI community: using limited hardware (a single RTX 3090) and open-source models to take on seemingly impossible tasks — and making the entire process fully transparent.
From a practical standpoint, rather than expecting a model to "prove the Riemann Hypothesis," it's more useful to treat these kinds of long-duration autonomous run experiments as test platforms for evaluating models' autonomous reasoning capabilities. Choosing a problem of moderate scope with a verifiable answer — rather than a century-defining challenge like the Riemann Hypothesis — would likely yield more actionable results.
Whether or not any agent ever solves a Millennium Prize Problem, experiments like this push us toward a better understanding of what happens when you "let a model run free" for dozens of hours — how far it can go, and where it ultimately stops.
Related articles

The Open Source Dilemma: A Non-Autoregressive Architecture Pioneer Overshadowed by Frontier Labs
An indie developer claims a frontier lab repackaged his year-old open-source non-autoregressive RL architecture as a breakthrough. We compare PPO sequence embeddings vs. RLCD parallel sampling and examine open source attribution gaps.

AI Plans an Entire Vineyard: A Real-World Experiment with 100 Grapevines
A Spokane hobbyist let Muse AI plan his entire vineyard — variety, spacing, irrigation, even the logo. He planted 100 Cabernet Franc vines and is documenting everything publicly.

Iceland's Treble Raises $18M to Bet on Voice Simulation Platform
Iceland-based voice simulation company Treble raises $18M. Its platform serves voice AI developers, AI wearables, and robotics firms. A deep dive into the technology and what the funding signals.