ElBruno.LocalLLMs.Rag 0.22.0

dotnet add package ElBruno.LocalLLMs.Rag --version 0.22.0
                    
NuGet\Install-Package ElBruno.LocalLLMs.Rag -Version 0.22.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="ElBruno.LocalLLMs.Rag" Version="0.22.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="ElBruno.LocalLLMs.Rag" Version="0.22.0" />
                    
Directory.Packages.props
<PackageReference Include="ElBruno.LocalLLMs.Rag" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add ElBruno.LocalLLMs.Rag --version 0.22.0
                    
#r "nuget: ElBruno.LocalLLMs.Rag, 0.22.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package ElBruno.LocalLLMs.Rag@0.22.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=ElBruno.LocalLLMs.Rag&version=0.22.0
                    
Install as a Cake Addin
#tool nuget:?package=ElBruno.LocalLLMs.Rag&version=0.22.0
                    
Install as a Cake Tool

ElBruno.LocalLLMs

NuGet NuGet Downloads Build Status License: MIT HuggingFace .NET GitHub stars Twitter Follow

Run local LLMs in .NET through IChatClient ๐Ÿง 

Run local LLMs in .NET through IChatClient โ€” the same interface you'd use for Azure OpenAI, Ollama, or any other provider. Powered by ONNX Runtime GenAI and BitNet.

What's New

The last 5 notable additions to the library. Updated with each NuGet release.

  • ๐ŸŽฏ ElBruno.LocalLLMs.Decisions โ€” new package for local System One decision models. Typed choice, score and yes/no answers with full probability distributions in a single forward pass (~600 ms on CPU), running Laya fully in-process on ONNX Runtime โ€” no server, no Python. Built for routing, triage, moderation and intent detection, where a chat model is slow and overkill. Call services.AddLocalDecisions() and inject IDecisionClient. See the Local Decisions Guide and the LocalDecisions sample.
  • โฌ†๏ธ .NET 10 only โ€” every project now single-targets net10.0. .NET 8 reaches end of support in November 2026, so multi-targeting was dropped ahead of it. Breaking: consumers still on .NET 8 must stay on v0.21.0 or upgrade to .NET 10.
  • ๐Ÿงฉ ElBruno.LocalLLMs.BlazorComponents โ€” new Razor Class Library with 7 ready-to-use Blazor components: ModelStatusCard (download progress bar + actions), ModelGallery (filterable grid), ModelSelector (two-way-bindable dropdown), ChatBox (streaming token display), EnvironmentDashboard (CPU/CUDA/DirectML badges), LocalLLMHealthBadge (nav-bar status dot), and RagPlayground. Call services.AddLocalLLMsBlazorComponents() to register. See the Blazor Components Guide and the BlazorDemo sample.
  • ๐Ÿง  GPT-OSS 20B support โ€” OpenAI's open-weight MoE model (Apache-2.0) now runs locally via the official onnxruntime/gpt-oss-20b-onnx artifacts. Adds the Harmony prompt format, channel-aware output filtering (chain-of-thought is stripped, never shown to users), Harmony tool calling, and a ReasoningEffort option. Two model IDs: gpt-oss-20b (CPU INT4) and gpt-oss-20b-cuda. See the GptOssChat sample. Also fixes a token-duplication bug that repeated the final token of every generation.
  • ๐Ÿš€ v0.20.12 โ€” Corrects sibling-package assembly versions, hardens vision token probing against model context limits, and verifies Fara smart image resizing for screenshot workflows.

Features

  • ๐Ÿงฉ Blazor components โ€” ModelStatusCard, ChatBox, ModelGallery, ModelSelector, EnvironmentDashboard, LocalLLMHealthBadge, RagPlayground via ElBruno.LocalLLMs.BlazorComponents (guide)
  • ๐ŸŽฏ Local decision models โ€” typed choice/score/yes-no answers with probabilities in one forward pass via ElBruno.LocalLLMs.Decisions (guide)
  • ๐Ÿ”Œ IChatClient implementation โ€” seamless integration with Microsoft.Extensions.AI
  • ๐Ÿ“ฆ Automatic model download โ€” models are fetched from HuggingFace on first use
  • ๐Ÿš€ Zero friction โ€” works out of the box with sensible defaults (Phi-3.5 mini)
  • ๐Ÿ–ฅ๏ธ Multi-hardware โ€” CPU, CUDA, and DirectML execution providers
  • ๐Ÿ’‰ DI-friendly โ€” register with AddLocalLLMs() or AddBitNetChatClient() in ASP.NET Core
  • ๐Ÿ”„ Streaming โ€” token-by-token streaming via GetStreamingResponseAsync
  • ๐Ÿ“Š Multi-model โ€” switch between Phi-3.5, Phi-4, Qwen2.5, Qwen3, Llama 3.2, MagenticBrain, and more
  • ๐Ÿ‘๏ธ Vision-language models โ€” run Fara 1.5-9B image+text models via LocalVisionChatClient
  • ๐Ÿค– Agentic models โ€” Qwen3 / MagenticBrain support for multi-agent orchestration loops
  • ๐ŸŽฏ Fine-tuned models โ€” pre-trained Qwen2.5 variants for tool calling and RAG (guide)
  • โšก BitNet support โ€” run 1.58-bit ternary models via bitnet.cpp with extreme efficiency (guide)
  • ๐Ÿ“ˆ OpenTelemetry diagnostics โ€” lifecycle activities and metrics for queued, first-token, completion, cancellation, and failure (guide)

Packages

Package NuGet Downloads Description
ElBruno.LocalLLMs NuGet Downloads Core library โ€” ONNX Runtime GenAI models via IChatClient
ElBruno.LocalLLMs.Rag NuGet Downloads RAG pipeline โ€” document chunking, indexing, retrieval
ElBruno.LocalLLMs.BitNet NuGet Downloads BitNet 1.58-bit models via bitnet.cpp + IChatClient
ElBruno.LocalLLMs.BlazorComponents NuGet Downloads Blazor components โ€” ModelStatusCard, ChatBox, ModelGallery, and more
ElBruno.LocalLLMs.Decisions NuGet Downloads System One decision models โ€” typed choice, score and yes/no answers

Installation

dotnet add package ElBruno.LocalLLMs

For CPU scenarios, no extra package is required โ€” the transitive buildTransitive shim copies onnxruntime-genai.dll automatically on Windows.

Add a runtime package only when you want a specific GPU provider:

# ๐ŸŸข NVIDIA GPU (CUDA):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.Cuda

# ๐Ÿ”ต Any Windows GPU โ€” AMD, Intel, NVIDIA (DirectML):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.DirectML

โš ๏ธ Add at most one GPU runtime package. Do not reference both Microsoft.ML.OnnxRuntimeGenAI.Cuda and Microsoft.ML.OnnxRuntimeGenAI.DirectML simultaneously.

If you use a GPU runtime package and want to disable the transitive CPU copy shim, set: <ElBrunoLocalLLMsDisableCpuNativeCopy>true</ElBrunoLocalLLMsDisableCpuNativeCopy> in your application .csproj.

๐Ÿš€ The library defaults to ExecutionProvider.Auto โ€” on Windows it tries DirectML โ†’ CUDA โ†’ CPU, and on Linux it tries CUDA โ†’ CPU. No code changes needed.

Quick Start

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

// Create a local chat client (downloads Phi-3.5 mini on first run)
using var client = await LocalChatClient.CreateAsync();

var response = await client.GetResponseAsync([
    new(ChatRole.User, "What is the capital of France?")
]);

Console.WriteLine(response.Text);

First Run

The first time you create a LocalChatClient, the model is downloaded from HuggingFace to your local cache directory (~2-4 GB). This typically takes 30-60 seconds depending on your internet connection.

Track download progress:

using var client = await LocalChatClient.CreateAsync(
    new LocalLLMsOptions { Model = KnownModels.Phi35MiniInstruct },
    progress: new Progress<ModelDownloadProgress>(p =>
    {
        var percent = (p.BytesDownloaded * 100) / p.TotalBytes;
        Console.WriteLine($"{p.FileName}: {percent:F1}%");
    })
);

Subsequent runs load instantly from cache (%LOCALAPPDATA%/ElBruno/LocalLLMs/models).

Skip auto-download if using a pre-downloaded model:

var options = new LocalLLMsOptions
{
    Model = KnownModels.Phi35MiniInstruct,
    ModelPath = "/path/to/local/model",
    EnsureModelDownloaded = false
};
using var client = await LocalChatClient.CreateAsync(options);

Streaming

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

using var client = await LocalChatClient.CreateAsync(new LocalLLMsOptions
{
    Model = KnownModels.Phi35MiniInstruct
});

await foreach (var update in client.GetStreamingResponseAsync([
    new(ChatRole.System, "You are a helpful assistant."),
    new(ChatRole.User, "Explain quantum computing in simple terms.")
]))
{
    Console.Write(update.Text);
}

GPU Acceleration

By default, ExecutionProvider.Auto tries GPU first and falls back to CPU automatically:

// Use explicit GPU provider (fails if CUDA not installed; use Auto to fallback to CPU)
var options = new LocalLLMsOptions
{
    ExecutionProvider = ExecutionProvider.Cuda
};

// Multi-GPU systems: select device ID
var options2 = new LocalLLMsOptions
{
    ExecutionProvider = ExecutionProvider.Cuda,
    GpuDeviceId = 1  // Use second GPU
};

Auto fallback behavior:

  • Windows + DirectML available โ†’ uses a Windows GPU through DirectML
  • Windows + DirectML unavailable, CUDA available โ†’ uses NVIDIA GPU
  • Linux + CUDA available โ†’ uses NVIDIA GPU
  • GPU unavailable โ†’ falls back to CPU (no errors, just slower)

โš ๏ธ CUDA note: ONNX Runtime GenAI 0.15.x expects CUDA 13.*, cuDNN 9.*, and the latest Microsoft Visual C++ 2015-2022 runtime. When those native libraries are missing, provider diagnostics now surface the exact DLL mismatch or missing dependency instead of entering the failing native path.

See Troubleshooting: GPU Setup for debugging GPU issues.

Model Metadata

Inspect model capabilities at runtime โ€” context window size, model name, and vocabulary:

using var client = await LocalChatClient.CreateAsync();

var metadata = client.ModelInfo;
Console.WriteLine($"Model:          {metadata?.ModelName}");
Console.WriteLine($"Context window: {metadata?.MaxSequenceLength}");
Console.WriteLine($"Vocab size:     {metadata?.VocabSize}");

This is useful for prompt-length validation, adaptive chunking, and model selection logic.

Dependency Injection

builder.Services.AddLocalLLMs(options =>
{
    options.Model = KnownModels.Phi35MiniInstruct;
    options.ExecutionProvider = ExecutionProvider.DirectML;
});

// Inject IChatClient anywhere
public class MyService(IChatClient chatClient) { ... }

Error Handling

The library provides structured exception types for graceful error handling:

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

try
{
    using var client = await LocalChatClient.CreateAsync();
    var response = await client.GetResponseAsync([
        new(ChatRole.User, "Your question here")
    ]);
}
catch (ExecutionProviderException ex)
{
    // GPU/provider-specific error (no CUDA, DirectML not available, etc.)
    Console.WriteLine($"Provider error: {ex.Message}");
}
catch (ModelCapacityExceededException ex)
{
    // Prompt/response too long for model's context window
    Console.WriteLine($"Capacity error: {ex.Message}");
    // Solution: use a larger model or truncate the prompt
}
catch (InvalidOperationException ex)
{
    // General operation error (model not found, download failed, etc.)
    Console.WriteLine($"Operation error: {ex.Message}");
}

Observability

LocalChatClient emits generation lifecycle diagnostics through ActivitySource and Meter, both named ElBruno.LocalLLMs.

using ElBruno.LocalLLMs.Diagnostics;

builder.Services.AddOpenTelemetry()
    .WithTracing(tracing => tracing.AddSource(LocalLLMsInstrumentation.ActivitySourceName))
    .WithMetrics(metrics => metrics.AddMeter(LocalLLMsInstrumentation.MeterName));

By default, telemetry excludes prompt and completion text. Opt in only when you want content attached:

var options = new LocalLLMsOptions
{
    CaptureTelemetryContent = true
};

See docs/observability.md for the lifecycle event contract, metric names, and Aspire wiring notes, and docs/cancellation.md for voice barge-in cancellation behavior.

Local Decisions

Sometimes you don't need prose โ€” you need a decision. ElBruno.LocalLLMs.Decisions runs a System One model that returns a typed answer with its full probability distribution in a single forward pass, fast enough to call on every request.

dotnet add package ElBruno.LocalLLMs.Decisions

The model runs in-process on ONNX Runtime โ€” no server to start and no Python. The weights are downloaded from elbruno/laya-onnx on first use (about 800 MB) and cached afterwards.

Ask several questions at once โ€” they share one forward pass, so four cost about what one costs:

using var client = new LayaOnnxDecisionClient();

DecisionResult result = await client.EvaluateAsync(
    new DecisionRequest("My invoice charged me twice and I want my money back.")
        .Choose("team", new Dictionary<string, string?>
        {
            ["billing"]   = "Payment, invoice and refund problems",
            ["technical"] = "Bugs, outages and API errors",
            ["sales"]     = "Pricing, upgrades and new purchases"
        })
        .Score("urgency", new[] { "no rush", "normal", "urgent", "critical" })
        .Ask("refund", "The customer is asking for a refund."));

ChoiceResult team = result.Choice("team");

// Gate on the probability so ambiguous tickets reach a human
// instead of being confidently misrouted.
string route = team.ChoiceOrNull(0.6) ?? "human-review";

Console.WriteLine(route);                                  // billing
Console.WriteLine(result.Score("urgency").Score);          // 1.67 of 3
Console.WriteLine(result.Probability("refund").IsTrue);    // True

โš ๏ธ Laya's public checkpoints report uncalibrated confidence. Fit DecisionThreshold against your own labelled examples before relying on a boolean verdict. Answers whose temperature bucket the checkpoint got wrong are flagged with CalibrationClamped โ€” see the Local Decisions Guide and the DecisionCalibration sample.

Cache Management

Inspect and manage the local model cache programmatically:

// Remove a model from the cache (no-op if not cached)
await LocalChatClient.DeleteModelFromCacheAsync(KnownModels.Phi35MiniInstruct);

// Or use a custom cache directory
await LocalChatClient.DeleteModelFromCacheAsync(
    KnownModels.Phi35MiniInstruct,
    cacheDirectory: @"D:\my-models");

// Get cached size in bytes for one model (0 if not downloaded)
long bytes = LocalChatClient.GetModelCacheSize(KnownModels.Phi35MiniInstruct);
Console.WriteLine($"Cached: {bytes / 1024 / 1024:N0} MB");

// List all cached models with size and last-modified date
var cached = LocalChatClient.ListCachedModels();
foreach (var repo in cached)
    Console.WriteLine($"{repo.LocalDirectory}  {repo.TotalSizeBytes / 1024 / 1024:N0} MB  {repo.LastModified:yyyy-MM-dd}");

// Same APIs available on LocalVisionChatClient for vision models
await LocalVisionChatClient.DeleteModelFromCacheAsync(KnownModels.Fara15_9B);
long visionBytes = LocalVisionChatClient.GetModelCacheSize(KnownModels.Fara15_9B);

The default cache directory is %LOCALAPPDATA%/ElBruno/LocalLLMs/models (Windows) or ~/.local/share/ElBruno/LocalLLMs/models (Linux/macOS).

These operations delegate to ElBruno.HuggingFace.Downloader which provides the underlying DeleteCachedFilesAsync, GetCachedSize, and ListCachedRepos implementation.

Troubleshooting

GPU not working? Use ExecutionProvider.Cpu explicitly. See GPU Setup Validation.

Out of memory? Try a smaller model:

var options = new LocalLLMsOptions
{
    Model = KnownModels.Qwen25_05BInstruct  // 0.5B instead of 3.8B
};

Model download fails?

  • Check your internet connection
  • For private HuggingFace models, set the HF_TOKEN environment variable

For detailed troubleshooting, see docs/troubleshooting-guide.md.

Supported Models

Tier Model Parameters ONNX ID
โšช Tiny TinyLlama-1.1B-Chat 1.1B โœ… Native tinyllama-1.1b-chat
โšช Tiny SmolLM2-1.7B-Instruct 1.7B โœ… Native smollm2-1.7b-instruct
โšช Tiny Qwen2.5-0.5B-Instruct 0.5B โœ… Native qwen2.5-0.5b-instruct
โšช Tiny Qwen2.5-1.5B-Instruct 1.5B โœ… Native qwen2.5-1.5b-instruct
โšช Tiny Gemma-2B-IT 2B โœ… Native gemma-2b-it
โšช Tiny Gemma-4-E2B-IT 5.1B (2B active) ๐Ÿ”„ Convert gemma-4-e2b-it
โšช Tiny StableLM-2-1.6B-Chat 1.6B โœ… Native stablelm-2-1.6b-chat
๐ŸŸข Small Phi-3.5 mini instruct 3.8B โœ… Native phi-3.5-mini-instruct
๐ŸŸข Small Qwen2.5-3B-Instruct 3B โœ… Native qwen2.5-3b-instruct
๐ŸŸข Small Llama-3.2-3B-Instruct 3B โœ… Native llama-3.2-3b-instruct
๐ŸŸข Small Gemma-2-2B-IT 2B โœ… Native gemma-2-2b-it
๐ŸŸข Small Gemma-4-E4B-IT 8B (4B active) ๐Ÿ”„ Convert gemma-4-e4b-it
๐ŸŸก Medium Qwen2.5-7B-Instruct 7B โœ… Native qwen2.5-7b-instruct
๐ŸŸก Medium Qwen2.5-Coder-7B-Instruct 7B โœ… Native qwen2.5-coder-7b-instruct
๐ŸŸก Medium Llama-3.1-8B-Instruct 8B โœ… Native llama-3.1-8b-instruct
๐ŸŸก Medium Mistral-7B-Instruct-v0.3 7B โœ… Native mistral-7b-instruct-v0.3
๐ŸŸก Medium Gemma-2-9B-IT 9B โœ… Native gemma-2-9b-it
๐ŸŸก Medium Gemma-4-12B-IT 12B ๐Ÿ”„ Convert gemma-4-12b-it
๐ŸŸก Medium Phi-4 14B โœ… Native phi-4
๐ŸŸก Medium DeepSeek-R1-Distill-Qwen-14B 14B โœ… Native deepseek-r1-distill-qwen-14b
๐ŸŸก Medium Mistral-Small-24B-Instruct 24B โœ… Native mistral-small-24b-instruct
๐Ÿ”ด Large Qwen2.5-14B-Instruct 14B โœ… Native qwen2.5-14b-instruct
๐Ÿ”ด Large Qwen2.5-32B-Instruct 32B โœ… Native qwen2.5-32b-instruct
๐Ÿ”ด Large Llama-3.3-70B-Instruct 70B โœ… ONNX llama-3.3-70b-instruct
๐Ÿ”ด Large Mixtral-8x7B-Instruct-v0.1 8x7B โœ… Native mixtral-8x7b-instruct-v0.1
๐Ÿ”ด Large DeepSeek-R1-Distill-Llama-70B 70B โœ… Native deepseek-r1-distill-llama-70b
๐Ÿ”ด Large Command-R (35B) 35B โœ… Native command-r-35b
๐Ÿ”ด Large Gemma-4-26B-A4B-IT 25.2B (3.8B active) ๐Ÿ”„ Convert gemma-4-26b-a4b-it
๐Ÿ”ด Large Gemma-4-31B-IT 30.7B ๐Ÿ”„ Convert gemma-4-31b-it
๐ŸŸฃ Next-Gen Qwen3-14B-Instruct 14.77B โœ… Native qwen3-14b-instruct
๐Ÿง  GPT-OSS GPT-OSS 20B (CPU INT4) 21B (3.6B active, MoE) โœ… Native gpt-oss-20b
๐Ÿง  GPT-OSS GPT-OSS 20B (CUDA INT4) 21B (3.6B active, MoE) โœ… Native gpt-oss-20b-cuda
๐Ÿค– Agentic MagenticBrain ~14.77B โœ… Native magentic-brain
๐Ÿ‘๏ธ VLM Fara 1.5-9B ~9.4B โœ… Native fara-1.5-9b

๐Ÿ”„ Convert = Use the conversion scripts in scripts/ to export ONNX locally before running the model.

ยน MagenticBrain ONNX: Native ONNX hosted at elbruno/MagenticBrain-onnx (INT4 quantized). Auto-downloads when EnsureModelDownloaded=true.

ยฒ Fara 1.5-9B ONNX: elbruno/Fara1.5-9B-onnx now includes the validated multimodal package (qwen3vl-vision.onnx, qwen3vl-embedding.onnx, patched genai_config.json, and ORT-compatible processor_config.json). See ONNX Conversion โ€” Fara.

ยณ GPT-OSS 20B: Apache-2.0, from the official onnxruntime/gpt-oss-20b-onnx repository. The CPU INT4 variant is a ~12 GB download, and because GPT-OSS is a mixture-of-experts model, CPU inference is slow โ€” prefer gpt-oss-20b-cuda with Microsoft.ML.OnnxRuntimeGenAI.Cuda when a GPU is available. GPT-OSS reasons before answering; that chain-of-thought is filtered out and never surfaced, per the model card. Reasoning depth is controlled by LocalLLMsOptions.ReasoningEffort.

Fine-Tuned Models

Pre-trained variants optimized for specific tasks. A fine-tuned 0.5B model often matches or exceeds a base 1.5B on its specialized task.

Model Size Task HuggingFace ID
Qwen2.5-0.5B-ToolCalling ~1 GB Tool/function calling elbruno/Qwen2.5-0.5B-LocalLLMs-ToolCalling
Qwen2.5-0.5B-RAG ~1 GB RAG with citations elbruno/Qwen2.5-0.5B-LocalLLMs-RAG
Qwen2.5-0.5B-Instruct ~1 GB General-purpose elbruno/Qwen2.5-0.5B-LocalLLMs-Instruct

See the Supported Models Guide for detailed model cards, performance benchmarks, and selection guidance.

Samples

Sample Description
HelloChat Minimal console chat
StreamingChat Token-by-token streaming
MultiModelChat Switch models at runtime
DependencyInjection ASP.NET Core DI registration
ToolCallingAgent Function calling and tool use
FineTunedToolCalling Fine-tuned model for improved tool calling
RagChatbot RAG pipeline with document retrieval
ZeroCloudRag Zero-cloud RAG pipeline with real local embeddings and LLM inference
BitNetChat BitNet 1.58-bit model chat completion
BitNetPerformance Performance benchmark: BitNet vs ONNX models
MagenticBrainAgent Multi-agent orchestration loop using Qwen3/MagenticBrain
FaraVisionAgent Vision-language model (Fara 1.5-9B) image+text inference
GptOssChat GPT-OSS 20B chat, streaming, reasoning effort, and tool calling
MagenticUIServer ASP.NET Core + SignalR multi-agent server (FileSurfer, WebFetcher, Coder)
ConsoleAppDemo Interactive console application
LocalDecisions Support-ticket triage with a local System One decision model
DecisionCalibration Calibration clamping and fp16 batch variance, made visible

๐ŸŒ Reference App: ElBruno.MagenticUI โ€” full Blazor Server port of microsoft/magentic-ui running locally with this library.

Requirements

  • .NET 10.0
  • CPU (default), NVIDIA GPU (CUDA), or Windows GPU (DirectML)
  • ~2-8 GB disk space per model (depending on size and quantization)

Building from Source

git clone https://github.com/elbruno/ElBruno.LocalLLMs.git
cd ElBruno.LocalLLMs
dotnet restore ElBruno.LocalLLMs.slnx
dotnet build ElBruno.LocalLLMs.slnx
dotnet test ElBruno.LocalLLMs.slnx

Run integration tests (downloads real models โ€” requires internet):

RUN_INTEGRATION_TESTS=true dotnet test ElBruno.LocalLLMs.slnx

Integration tests validate the full lifecycle (download โ†’ infer โ†’ cache hit โ†’ delete) for all 35 supported models. See docs/tests/README.md for details.

Documentation

๐Ÿค Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

๐Ÿ“„ License

This project is licensed under the MIT License โ€” see the LICENSE file for details.

๐Ÿ‘‹ About the Author

Hi! I'm ElBruno ๐Ÿงก, a passionate developer and content creator exploring AI, .NET, and modern development practices.

Made with โค๏ธ by ElBruno

If you like this project, consider following my work across platforms:

  • ๐Ÿ“ป Podcast: No Tienen Nombre โ€” Spanish-language episodes on AI, development, and tech culture
  • ๐Ÿ’ป Blog: ElBruno.com โ€” Deep dives on embeddings, RAG, .NET, and local AI
  • ๐Ÿ“บ YouTube: youtube.com/elbruno โ€” Demos, tutorials, and live coding
  • ๐Ÿ”— LinkedIn: @elbruno โ€” Professional updates and insights
  • ๐• Twitter: @elbruno โ€” Quick tips, releases, and tech news

๐Ÿ™ Acknowledgments

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (1)

Showing the top 1 NuGet packages that depend on ElBruno.LocalLLMs.Rag:

Package Downloads
ElBruno.LocalLLMs.BlazorComponents

Blazor UI components for ElBruno.LocalLLMs โ€” model status cards, chat boxes, download management, environment diagnostics, RAG playground and more.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.22.0 7 9/30/2026
0.21.0 169 8/11/2026
0.20.12 111 8/11/2026
0.20.11 122 8/11/2026
0.20.10 128 8/10/2026
0.20.9 135 8/2/2026
0.20.8 149 8/1/2026
0.20.7 121 8/1/2026
0.20.6 127 8/1/2026
0.20.5 113 7/30/2026
0.20.4 115 7/29/2026
0.20.3 114 7/28/2026
0.20.1 113 7/24/2026
0.20.0 109 7/23/2026
0.19.1 134 7/3/2026
0.19.0 131 7/3/2026
0.18.0 136 6/5/2026
0.17.0 119 6/3/2026
0.16.0 127 4/17/2026
0.15.0 119 4/16/2026
Loading failed