Developers can now run the Qwen3-0.6B language model entirely within an Internet Computer canister following the release of llama_cpp_canister v0.13.0, allowing AI applications to perform verifiable inference directly on-chain.
The latest version enables users to deploy models built with the ggml.org llama.cpp framework as smart contracts on the Internet Computer. The release highlights Qwen3-0.6B as the recommended default model, supporting multi-turn conversations with a context window of around 12,000 words.
Unlike traditional AI systems that rely on external servers for processing, llama_cpp_canister is designed to keep inference inside a canister environment. This allows developers to build AI agents where model execution can be verified within the same decentralised infrastructure.
The update supports running large language models in the GGUF format, giving developers the option to deploy different models depending on their application requirements. The team has recommended Qwen3-0.6B in non-thinking mode for conversational use cases requiring extended back-and-forth interactions.
Alongside on-chain model execution, llama_cpp_canister v0.13.0 includes developer-focused improvements such as open-source availability under the MIT licence, documentation resources and automated quality checks through continuous integration and deployment workflows.
The project also includes tools designed to simplify development, testing and deployment, including a smoke-testing framework using pytest. These features are intended to help developers experiment with AI applications while maintaining a clear development process.
Running AI models directly within smart contracts remains an emerging area, with developers exploring ways to combine artificial intelligence with decentralised infrastructure. Challenges such as computing requirements, model size and performance optimisation continue to influence how these systems are built and adopted.
With llama_cpp_canister, developers on the Internet Computer can experiment with verifiable AI workloads that operate within the network itself, opening opportunities for applications that require both autonomous AI capabilities and transparent execution.
Dear Reader,
Ledger Life is an independent platform dedicated to covering the Internet Computer (ICP) ecosystem and beyond. We focus on real stories, builder updates, project launches, and the quiet innovations that often get missed.
We’re not backed by sponsors. We rely on readers like you.
If you find value in what we publish—whether it’s deep dives into dApps, explainers on decentralised tech, or just keeping track of what’s moving in Web3—please consider making a donation. It helps us cover costs, stay consistent, and remain truly independent.
Your support goes a long way.
🧠 ICP Principal: ins6i-d53ug-zxmgh-qvum3-r3pvl-ufcvu-bdyon-ovzdy-d26k3-lgq2v-3qe
🧾 ICP Address: f8deb966878f8b83204b251d5d799e0345ea72b8e62e8cf9da8d8830e1b3b05f
Every contribution helps keep the lights on, the stories flowing, and the crypto clutter out.
Thank you for reading, sharing, and being part of this experiment in decentralised media.
—Team Ledger Life





Community Discussion