Native .NET AI inference engine for GGUF models — text, reasoning, multimodal input, embeddings, image generation and editing, and video with audio. Run it from the CLI, browser chat, Ollama/OpenAI-compatible APIs, or TensorAgent, the local app for iPhone, iPad, Mac and Windows. The .NET runtime offers managed CPU and native accelerator backends; published comparisons use identical GGUF files and hardware. The optional TensorSharp.AgentHost layer adds Agent Skills, a bounded, in-process model-to-tool loop for file and shell work, and bounded automatic subagent delegation.
- Local, native .NET inference. Run GGUF text and multimodal models from the CLI, browser UI, or Ollama/OpenAI-compatible APIs.
- Broad model and media support. Current source covers modern text models, DiffusionGemma text diffusion, vision/audio input, PDF, Qwen-Image-2.1 image generation and editing with masks and LoRA plug-ins, MiniMax-H3 video with native 32 kHz stereo audio, and Wan 2.1/2.2 video. See the model cards.
- Text and code embeddings. GGUF BERT/XLM-R encoders with OpenAI/Ollama batch embedding APIs for Snowflake Arctic Embed and MiniLM; see the embedding guide.
- Measured performance. TensorSharp is benchmarked against
llama.cppon identical models and hardware. Results are specific to the measured model, backend, and workload. See Benchmarks. - Agentic work.
TensorSharp.AgentHostadds bounded Agent Skills, code tools, and automatic subagent delegation with independent contexts, private workspaces, dependency scheduling, and read-only defaults. - TensorAgent for phones and desktops. One app for local chat, multimodal input, code and document work, image generation/editing, and short video with audio. Its eleven-model catalog is gated by device memory; image/video models need desktop memory tiers. The interface supports English, Simplified and Traditional Chinese, Japanese, Korean, Spanish, French and German. See TensorAgent for source builds, platform differences and measured coverage.
- Production-friendly building blocks. Continuous batching and the paged, Radix prefix-shared KV cache are on by default; speculative decoding, tensor parallelism, and configurable security boundaries are available when you need them. See Features, Usage, and the current project status.
- Text, reasoning, and multimodal LLMs: DeepSeek V4 Flash / V4.1 Flash, GLM 5.x, Gemma 4, Qwen 3.5 / 3.6 / 3.8 27B, Qwen 3.8 Flash Next, Bonsai2 (Qwen family), GPT OSS, Nemotron-H, Mistral 3, Hunyuan Dense, and Muse-Glimmer.
- Text diffusion: DiffusionGemma, including Jev typed decision inference at
/v1/systemone, over text, images, uploaded documents, sampled video frames and audio transcripts (configured ASR companion). - Image generation/editing and video generation: Qwen-Image-2.1, MiniMax-H3 (video + stereo audio), and Wan 2.1 / 2.2.
- Text and code embeddings: BERT / XLM-R encoders — Snowflake Arctic Embed L v2.0 and all-MiniLM-L6-v2.
Backend, modality, feature support, and validation coverage vary by model. See the supported models tables, the model cards, and the embedding guide for details.
Recent source additions include Qwen-Image-2.1 masked edits with exact protected pixels and optional processing of the selected region, twelve TensorAgent LoRA plug-ins for speed, style and editing, and Qwen3.8 Flash Next on a 48 GB Mac using SSD-backed weights. Multi-GPU --layer-split and --tp are separate controls; support and performance depend on the architecture and quantization. These source features may be ahead of the published CLI/server packages; TensorAgent currently requires a source build.
| Qwen inference and agentic runtimes | Gemma 4 and multimodal inference |
|---|---|
![]() |
![]() |
| Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent | From Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B |
| Build Qwen dense/MoE inference and controlled agent workflows in C#. Follow tensors, tokenization, attention, expert routing, quantization, and caching through GPU acceleration, multimodal execution, tools, skills, sandboxed code, and desktop/mobile deployment with TensorSharp and TensorAgent. | Build a multimodal inference engine in C#/.NET with Gemma 4 E4B, from tensors, GGUF model loading, quantization, and tokenization to text, image, video, and audio execution. Connect correctness checks and serving optimizations to the running TensorSharp code. |
| Buy on Amazon | Buy on Amazon |
Explore both books and their repository reading paths
Prefer a prebuilt application? The Releases page provides self-contained CLI and Server archives for Windows x64 (CPU/CUDA), Linux x64 (CPU/CUDA), and macOS arm64.
To build from source you need the full .NET 10 SDK (how to install it), git, curl, CMake 3.20+, and the toolchain for your GPU. Then run the verified Gemma 4 E4B model (7.48 GiB). On Windows with an NVIDIA GPU (PowerShell):
git clone https://github.com/zhongkaifu/TensorSharp.git; Set-Location TensorSharp
New-Item -ItemType Directory -Force models | Out-Null
curl.exe -L --fail "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models\gemma-4-E4B-it-Q8_0.gguf
'Answer in one short sentence: what is TensorSharp?' | Set-Content prompt.txt
$env:TENSORSHARP_GGML_NATIVE_ENABLE_CUDA = 'ON'
dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models\gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cudaOn other machines, change the backend (see Pick a backend):
- macOS (Apple Silicon): drop the CUDA environment variable and use
--backend ggml_metal. - Linux + NVIDIA: prefix the
dotnet runwithTENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ONand use--backend ggml_cuda. - AMD / Intel / NVIDIA Vulkan: set
TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ONand use--backend ggml_vulkan.
Host the same model as a server: a browser chat at http://localhost:5000 plus Ollama- and OpenAI-compatible APIs.
dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 512The server listens on
0.0.0.0:5000with no built-in authentication or TLS; keep it behind a firewall or an authenticated HTTPS reverse proxy.
| Your hardware | Backend |
|---|---|
| Apple Silicon (Mac) | --backend ggml_metal |
| Windows / Linux + NVIDIA GPU | --backend ggml_cuda |
| Windows / Linux + AMD / Intel / NVIDIA GPU | --backend ggml_vulkan |
| No GPU | --backend ggml_cpu (native kernels), or --backend cpu (pure C#, no native dependencies) |
The Getting started guide has the rest: installing the SDK on each platform, multi-GPU and multi-node runs, NVIDIA DGX Spark, multimodal input, embeddings, and making it fast. Every option is in the CLI and Server references, and both programs print them with --help.
dotnet build TensorSharp.slnx also builds TensorAgent's available desktop heads and the iOS simulator head on Apple Silicon when the selected SDK has the required MAUI workloads and staged native/Python files. Missing prerequisites skip the affected app head with a warning; see TensorAgent build instructions.
One engine, four ways to use it, each an unedited capture of a real run.
What each run shows, step by step: Screenshots.
TensorSharp and llama.cpp run identical GGUF files on the same NVIDIA RTX 3080 Laptop GPU (16 GB), each on its GGML CUDA and Vulkan builds. Each number is TensorSharp's speedup over llama.cpp on the same backend (geomean, single-stream, greedy, MTP off); above 1.0× means TensorSharp is faster.
| Model | Backend | decode | prefill | TTFT |
|---|---|---|---|---|
| Gemma 4 E4B it (Q8_0, dense multimodal) | CUDA | 1.02× | 1.28× | 1.27× |
| Gemma 4 E4B it (Q8_0, dense multimodal) | Vulkan | 1.00× | 1.05× | 1.03× |
| Gemma 4 12B it (QAT UD-Q4_K_XL, dense) | CUDA | 1.04× | 1.17× | 1.16× |
| Gemma 4 12B it (QAT UD-Q4_K_XL, dense) | Vulkan | 1.21× | 1.04× | 1.03× |
| Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE) | CUDA | 0.98× | 1.28× | 1.27× |
| Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE) | Vulkan | 0.87× | 1.04× | 1.03× |
| Qwen 3.6 27B (UD-IQ2_XXS, dense) | CUDA | 1.07× | 0.96× | 0.95× |
| Qwen 3.6 27B (UD-IQ2_XXS, dense) | Vulkan | 1.02× | 0.85× | 0.84× |
What these numbers mean, how to rerun them, and the head-to-heads of models too large for this GPU: Benchmarks.
New here? The sections above are all you need to get running. Everything else is detailed reference:
| Doc | What's inside |
|---|---|
| TensorSharp and TensorAgent book guide | Building LLM Inference Engines and Agentic Runtimes from Scratch, plus From Tensors to Tokens: introductions, Amazon links, and repository reading paths |
| Getting started | The full first-run guide: the .NET SDK on each platform, every backend, multi-GPU and multi-node runs, NVIDIA DGX Spark, embeddings, choosing a backend, and making it fast |
| Supported models | Implemented model families and their validation scope: example GGUFs, modalities, thinking, tools, and speculative decoding |
| Benchmarks | TensorSharp against llama.cpp on the same GPU and files, and the head-to-heads of larger models |
| Screenshots | The CLI, the Web UI, and TensorAgent on iPhone and Mac at work, with what each run did |
| Model Downloads | Per-model huggingface-cli download + run quick reference (quant tiers, projectors, companions) |
| Usage | Full CLI reference (options, interactive REPL, JSONL batch), server hosting, logging, HTTP API examples, backends, and the env-var matrix |
| Features | Deep dives on continuous batching, speculative decoding, tool calling, thinking mode, multimodal, MoE, KV codecs, and more |
| Configuration files | Put options in a reusable JSON file with ${variables} and auto-downloading models |
| Development | Prerequisites, building the native GGML/MLX libraries, repository layout, package boundaries, internal architecture, and the test harness |
| Per-model architecture cards | End-to-end docs of each architecture (forward graph, components, parameters, prefill/decode optimizations) |
| Paged attention & continuous batching | The vLLM-style paged KV cache, prefix sharing, and iteration-level scheduler |
| Agent Skills & agentic work | The SKILL.md format, progressive disclosure and its budget, the in-process tool loop, sandboxed code execution, workspaces and artifacts, the path/ZIP/exec security model, and the HTTP + C# surfaces |
| Multiple agents | Automatic task delegation, private child workspaces, dependency scheduling, permission limits, server controls, and reproducible evaluation |
| Browser automation skill (Playwright) | Running the bundled playwright skill, which drives a browser through @playwright/cli via skills_run: the flags it needs, the macOS Chromium-sandbox config, account handoff, and TensorAgent desktop hosting (not iOS) |
| Speculative decoding | The three-layer design (model adapter / algorithm / speculator weights), the shipped auto / draft-head / block / ngram algorithms, and what to write to add a new one |
| Environment variable feature matrix | Which high-impact runtime flags affect which models, backends, and prompt types |
| Engine comparison report | Full per-scenario TensorSharp vs llama.cpp tables |
| ggml_metal vs llama.cpp | Head-to-head prefill/decode on Apple Silicon, the four graph-construction gaps it found, and what each was worth |
| Test/benchmark matrix runner | Sweep model × backend × feature × env-var cells and generate regression reports |
| Server API examples | Complete curl and Python examples for the server surface |
Actively developed, and the source tree runs ahead of the published packages.
| Area | Where it stands |
|---|---|
| Models | A dozen autoregressive families plus text diffusion, image generation and editing, and video with audio. See Supported models. |
| Inference hosts | CLI, Web UI, Ollama- and OpenAI-compatible APIs, and the TensorAgent app for iPhone, iPad, Mac and Windows. TensorAgent is source-only. |
| Backends | Pure C# CPU, direct CUDA/cuBLAS, MLX Metal, and GGML CPU/Metal/CUDA/Vulkan, with per-architecture exceptions. |
| Serving features | Continuous batching with a shared prefix cache, speculative decoding, tensor parallelism, structured output, and tool calling. |
| Agentic work | Agent Skills, sandboxed file and shell tools, and bounded sub-agents. See Agent Skills and Multiple agents. |
| TensorAgent | Eleven catalog entries, saved chats and artifacts, masked image edits and LoRA choices, eight interface languages, and persisted text-turn statistics. Media generation has been measured on a Mac; iOS media generation and Windows image/audio/video generation remain unverified. |
Per-area detail (which architecture runs on which backend, which features each family supports, and the known limits) is in the status matrix.
Zhongkai Fu
See LICENSE for details.






