Researchers working on the Internet Computer Protocol (ICP) say they have improved the efficiency of running an artificial intelligence language model directly onchain, allowing it to generate almost three times as many tokens within the same computing limits.
The work, carried out by researchers from Meotis, Kaizen Corp and ORIGYN, demonstrated that a small language model could perform inference entirely inside ICP canisters without relying on external cloud servers. According to the research team, the optimised model produced 29 tokens under the same instruction budget where the previous version generated around 10.
The findings have been published alongside the research paper, source code and performance measurements, allowing other developers to verify and reproduce the results.
While the work focuses on a relatively small language model rather than large-scale systems such as ChatGPT, researchers say it highlights the potential for AI applications and autonomous agents to operate directly on decentralised infrastructure.
Running AI inference onchain allows network participants to verify how models are executed through the protocol rather than depending on infrastructure controlled by a single cloud provider. Supporters argue this approach could improve transparency and trust for applications where verifiable execution is important.
The research has already been independently validated by members of the ICP developer community.
A developer known as “icpp” confirmed the performance gains after upgrading to the latest version of llama.cpp, reporting that the updated build generated around 28 tokens per update call using the same Qwen2.5-0.5B-Instruct Q8_0 model. The results were measured using the same methodology and verified on the ICP mainnet.
The collaboration also led to improvements in the underlying development tools. The latest release of icpp-pro 5.4.2 addresses an environment handling issue at the toolchain level, simplifying future upgrades of llama.cpp by removing the need for application-specific workarounds.
Shortly afterwards, llama_cpp_canister v0.12.0 was released, incorporating the latest llama.cpp upgrade and restoring the higher throughput achieved during testing. The project credits the Meotis, Kaizen Corp and ORIGYN researchers for identifying the missing WebAssembly SIMD optimisation and contributing an approach that has now been integrated into the software.
Members of the ICP community welcomed the development. One contributor, posting under the name Future_Proof, said onchain embedding and inference could enable a range of new AI applications.
The research adds to ongoing efforts across the blockchain and decentralised computing sectors to move AI workloads closer to the execution layer of distributed networks. Most large language models today continue to rely on conventional cloud infrastructure because of the substantial computing resources they require.
Researchers involved in the ICP project acknowledge that the latest work represents an incremental advance rather than a replacement for cloud-hosted frontier models. The experiments used a compact open-source model designed to fit within the network’s current execution limits.
Even so, the nearly threefold improvement suggests optimisation of software, compilers and execution environments can materially improve the performance of on-chain AI. As research continues, developers are expected to explore larger models, improved inference techniques and new applications where verifiable AI execution offers advantages over traditional cloud-based deployments.
Dear Reader,
Ledger Life is an independent platform dedicated to covering the Internet Computer (ICP) ecosystem and beyond. We focus on real stories, builder updates, project launches, and the quiet innovations that often get missed.
We’re not backed by sponsors. We rely on readers like you.
If you find value in what we publish—whether it’s deep dives into dApps, explainers on decentralised tech, or just keeping track of what’s moving in Web3—please consider making a donation. It helps us cover costs, stay consistent, and remain truly independent.
Your support goes a long way.
🧠 ICP Principal: ins6i-d53ug-zxmgh-qvum3-r3pvl-ufcvu-bdyon-ovzdy-d26k3-lgq2v-3qe
🧾 ICP Address: f8deb966878f8b83204b251d5d799e0345ea72b8e62e8cf9da8d8830e1b3b05f
Every contribution helps keep the lights on, the stories flowing, and the crypto clutter out.
Thank you for reading, sharing, and being part of this experiment in decentralised media.
—Team Ledger Life





Community Discussion