llama_cpp_canister brings Qwen3 AI model fully on-chain to the Internet Computer

Developers can now run the Qwen3-0.6B language model entirely within an Internet Computer canister following the release of llama_cpp_canister v0.13.0, allowing AI applications to perform verifiable inference directly on-chain.

The latest version enables users to deploy models built with the ggml.org llama.cpp framework as smart contracts on the Internet Computer. The release highlights Qwen3-0.6B as the recommended default model, supporting multi-turn conversations with a context window of around 12,000 words.

Unlike traditional AI systems that rely on external servers for processing, llama_cpp_canister is designed to keep inference inside a canister environment. This allows developers to build AI agents where model execution can be verified within the same decentralised infrastructure.

The update supports running large language models in the GGUF format, giving developers the option to deploy different models depending on their application requirements. The team has recommended Qwen3-0.6B in non-thinking mode for conversational use cases requiring extended back-and-forth interactions.

Alongside on-chain model execution, llama_cpp_canister v0.13.0 includes developer-focused improvements such as open-source availability under the MIT licence, documentation resources and automated quality checks through continuous integration and deployment workflows.

The project also includes tools designed to simplify development, testing and deployment, including a smoke-testing framework using pytest. These features are intended to help developers experiment with AI applications while maintaining a clear development process.

Running AI models directly within smart contracts remains an emerging area, with developers exploring ways to combine artificial intelligence with decentralised infrastructure. Challenges such as computing requirements, model size and performance optimisation continue to influence how these systems are built and adopted.

With llama_cpp_canister, developers on the Internet Computer can experiment with verifiable AI workloads that operate within the network itself, opening opportunities for applications that require both autonomous AI capabilities and transparent execution.


Dear Reader,

Ledger Life is an independent platform dedicated to covering the Internet Computer (ICP) ecosystem and beyond. We focus on real stories, builder updates, project launches, and the quiet innovations that often get missed.

We’re not backed by sponsors. We rely on readers like you.

If you find value in what we publish—whether it’s deep dives into dApps, explainers on decentralised tech, or just keeping track of what’s moving in Web3—please consider making a donation. It helps us cover costs, stay consistent, and remain truly independent.

Your support goes a long way.

🧠 ICP Principal: ins6i-d53ug-zxmgh-qvum3-r3pvl-ufcvu-bdyon-ovzdy-d26k3-lgq2v-3qe

🧾 ICP Address: f8deb966878f8b83204b251d5d799e0345ea72b8e62e8cf9da8d8830e1b3b05f

Every contribution helps keep the lights on, the stories flowing, and the crypto clutter out.

Thank you for reading, sharing, and being part of this experiment in decentralised media.
—Team Ledger Life

0

Community Discussion

Loading discussion…

LEAVE A REPLY

Please enter your comment!
Please enter your name here

More like this

DFINITY announces final days of MULTI/DEX Season I with...

DFINITY founder Dominic Williams has announced that the first season of the MULTI/DEX trading competition will conclude...

Caffeine adds collaboration tools for shared app development

Caffeine has introduced new collaboration features that allow users to invite others to work together on app...

DFINITY executive explains ICP’s focus on sovereign cloud infrastructure

DFINITY chief business officer Pierre Samaties has outlined the broader vision behind the Internet Computer Protocol (ICP),...