iFlytek Spark X2.5 4B On-Device Model Tested: Running Claude Code Offline to Complete Agent Tasks

iFlytek's 4B/1.7B on-device models deliver full Agent task capabilities in a completely offline, local environment.
iFlytek's open-source Spark X2.5 4B and 1.7B on-device models challenge the assumption that small models can only handle simple conversations. Real-world testing with Claude Code in a fully offline environment confirmed they can complete complex tasks — including multi-version contract diffing, cross-language review analysis, bulk local file archiving, and full code project review and repair — with all data remaining on-device. This satisfies both privacy requirements and eliminates cloud API costs and reliability concerns. As these small models gain Agent capabilities, million-token context windows, and tool-calling support, the benchmark for AI usefulness has shifted from parameter count to practical problem-solving — making on-device models a pragmatic choice for secure, capable local AI.
The Capability Ceiling of Small Models Is Being Raised
For a long time, the industry held a stereotype about small models: fewer parameters meant weaker performance — fine for simple conversations, but hopeless when faced with complex, multi-step tasks. A recent development in the open-source world is challenging that assumption.
iFlytek has open-sourced two on-device small models — Spark X2.5 4B and Spark X2.5 1.7B. The former is positioned as the workhorse capable of handling heavier agentic workloads, while the latter is lighter and designed to run on smartphones or edge devices. What's truly surprising isn't their parameter count, but the fact that they can reliably call tools and execute Agent tasks offline, locally, without a high-end GPU.
To put this to the test, the most direct approach was clear: load the 4B model into Claude Code, cut the internet connection entirely, and see if it can actually get work done.
Offline Testing: Contract Comparison and Multilingual Analysis
The first test tackled a classic pain point in daily office work — comparing multiple versions of a contract. The task required the model to compare the key clauses across three contracts, flag each change with its clause reference and page number, then generate a contract diff table and a list of unresolved questions.
The entire process was completed offline, with no data leaving the local machine. The model accurately identified the differences across all three contract versions, organized them into a summary table, and separately flagged conflicting language for confirmation. The results checked out.

The second test challenged the model's multilingual capabilities. Spark X2.5 4B claims support for over 200 languages, so a cross-border e-commerce scenario was set up: a batch of product reviews in English and Spanish needed to be analyzed for the most and least satisfying aspects, common issues tallied, a Chinese-language analysis report generated, and customer service reply templates created for both languages.
From cross-language comprehension and issue distribution stats, to identifying product strengths and pain points, to in-depth analysis of recurring complaints, and finally to generating response templates — the entire pipeline ran successfully. A 4B on-device model completing such a long task chain in one go genuinely exceeded many expectations.

Local Office Use: File Organization and Code Review
One of the greatest strengths of on-device small models is handling sensitive files that can't be uploaded to the cloud.
The third test scenario involved bulk organization of client files. The instruction was straightforward: rename documents in a folder according to a specified format, then move them all into a new archive folder. The model successfully identified the client name and creation date for each file, applied a consistent renaming rule in bulk, and automatically created the destination folder to complete the archiving.
What had been a scattered, inconsistently named pile of client documents on the desktop was neatly organized — no need to open each file individually, copy names, rename them, and move them one by one. And crucially, sensitive client data never left the machine.

The fourth test raised the bar considerably: reviewing a real code project while completely offline. A local folder containing a code file with multiple issues was provided, and Spark X2.5 4B was asked to read and audit it directly.
The model independently read through the project, reviewed the code step by step, compiled a summary of all bugs, and provided a specific description and fix for each one. When errors were encountered, it continued investigating and applying corrections until the project ran again. This wasn't simply auto-completing a snippet of code — it completed the full loop of reading the project, identifying problems, making fixes, running tests, and verifying results, entirely on the local machine, with code and business data never leaving the device.

Worth noting: both models can be deeply integrated with Claude Code, Codex, and Open Cloud, and were trained entirely on domestically produced computing infrastructure.
Why On-Device Small Models Deserve Attention
Some might ask: cloud-based large models are more powerful, so why bother with small models?
The answer is that not every task is suited to sending data offsite. Many companies have strict policies against data leaving the premises; cloud-based Agents can be unreliable; and the cost of frequent API calls quickly becomes difficult to control.
In this context, if a 1.7B or 4B small model can call tools, execute multi-step tasks, operate as an Agent, and handle a 1-million-token context window, then running it locally is no longer just an offline chat widget — it becomes a local AI assistant that can actually do real work. It satisfies both requirements: data stays secure, and real capability is available.
From a technical standpoint, iFlytek's approach is somewhat contrarian: while most players are scaling models up and pushing capabilities to the cloud, it chose to compress Agent-level capabilities into 1.7B and 4B models — letting ordinary computers and smartphones handle real workloads locally, even without an internet connection.
A Model's Value Is No Longer Just About Size
This real-world test delivers a more fundamental message: the true measure of whether AI is useful ultimately comes down to whether we can confidently hand off our actual work to it.
From that perspective, model size is no longer the only metric — what matters is whether it can solve real problems. When a local small model can handle genuine tasks and everyday challenges in an offline environment, that in itself is valuable enough.
Making AI more powerful is one kind of progress. Making that capability accessible and usable across more devices is equally important progress. For individuals and organizations who care about data privacy and want to reduce operating costs, the breakthrough represented by this new generation of on-device small models is well worth experiencing firsthand.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.