ElBruno.LocalLLMs.Rag
0.22.0
dotnet add package ElBruno.LocalLLMs.Rag --version 0.22.0
NuGet\Install-Package ElBruno.LocalLLMs.Rag -Version 0.22.0
<PackageReference Include="ElBruno.LocalLLMs.Rag" Version="0.22.0" />
<PackageVersion Include="ElBruno.LocalLLMs.Rag" Version="0.22.0" />
<PackageReference Include="ElBruno.LocalLLMs.Rag" />
paket add ElBruno.LocalLLMs.Rag --version 0.22.0
#r "nuget: ElBruno.LocalLLMs.Rag, 0.22.0"
#:package ElBruno.LocalLLMs.Rag@0.22.0
#addin nuget:?package=ElBruno.LocalLLMs.Rag&version=0.22.0
#tool nuget:?package=ElBruno.LocalLLMs.Rag&version=0.22.0
ElBruno.LocalLLMs
Run local LLMs in .NET through IChatClient ๐ง
Run local LLMs in .NET through IChatClient โ the same interface you'd use for Azure OpenAI, Ollama, or any other provider. Powered by ONNX Runtime GenAI and BitNet.
What's New
The last 5 notable additions to the library. Updated with each NuGet release.
- ๐ฏ
ElBruno.LocalLLMs.Decisionsโ new package for local System One decision models. Typedchoice,scoreand yes/no answers with full probability distributions in a single forward pass (~600 ms on CPU), running Laya fully in-process on ONNX Runtime โ no server, no Python. Built for routing, triage, moderation and intent detection, where a chat model is slow and overkill. Callservices.AddLocalDecisions()and injectIDecisionClient. See the Local Decisions Guide and the LocalDecisions sample. - โฌ๏ธ .NET 10 only โ every project now single-targets
net10.0. .NET 8 reaches end of support in November 2026, so multi-targeting was dropped ahead of it. Breaking: consumers still on .NET 8 must stay onv0.21.0or upgrade to .NET 10. - ๐งฉ
ElBruno.LocalLLMs.BlazorComponentsโ new Razor Class Library with 7 ready-to-use Blazor components:ModelStatusCard(download progress bar + actions),ModelGallery(filterable grid),ModelSelector(two-way-bindable dropdown),ChatBox(streaming token display),EnvironmentDashboard(CPU/CUDA/DirectML badges),LocalLLMHealthBadge(nav-bar status dot), andRagPlayground. Callservices.AddLocalLLMsBlazorComponents()to register. See the Blazor Components Guide and the BlazorDemo sample. - ๐ง GPT-OSS 20B support โ OpenAI's open-weight MoE model (Apache-2.0) now runs locally via the official
onnxruntime/gpt-oss-20b-onnxartifacts. Adds the Harmony prompt format, channel-aware output filtering (chain-of-thought is stripped, never shown to users), Harmony tool calling, and aReasoningEffortoption. Two model IDs:gpt-oss-20b(CPU INT4) andgpt-oss-20b-cuda. See the GptOssChat sample. Also fixes a token-duplication bug that repeated the final token of every generation. - ๐
v0.20.12โ Corrects sibling-package assembly versions, hardens vision token probing against model context limits, and verifies Fara smart image resizing for screenshot workflows.
Features
- ๐งฉ Blazor components โ
ModelStatusCard,ChatBox,ModelGallery,ModelSelector,EnvironmentDashboard,LocalLLMHealthBadge,RagPlaygroundviaElBruno.LocalLLMs.BlazorComponents(guide) - ๐ฏ Local decision models โ typed choice/score/yes-no answers with probabilities in one forward pass via
ElBruno.LocalLLMs.Decisions(guide) - ๐
IChatClientimplementation โ seamless integration with Microsoft.Extensions.AI - ๐ฆ Automatic model download โ models are fetched from HuggingFace on first use
- ๐ Zero friction โ works out of the box with sensible defaults (Phi-3.5 mini)
- ๐ฅ๏ธ Multi-hardware โ CPU, CUDA, and DirectML execution providers
- ๐ DI-friendly โ register with
AddLocalLLMs()orAddBitNetChatClient()in ASP.NET Core - ๐ Streaming โ token-by-token streaming via
GetStreamingResponseAsync - ๐ Multi-model โ switch between Phi-3.5, Phi-4, Qwen2.5, Qwen3, Llama 3.2, MagenticBrain, and more
- ๐๏ธ Vision-language models โ run Fara 1.5-9B image+text models via
LocalVisionChatClient - ๐ค Agentic models โ Qwen3 / MagenticBrain support for multi-agent orchestration loops
- ๐ฏ Fine-tuned models โ pre-trained Qwen2.5 variants for tool calling and RAG (guide)
- โก BitNet support โ run 1.58-bit ternary models via bitnet.cpp with extreme efficiency (guide)
- ๐ OpenTelemetry diagnostics โ lifecycle activities and metrics for queued, first-token, completion, cancellation, and failure (guide)
Packages
Installation
dotnet add package ElBruno.LocalLLMs
For CPU scenarios, no extra package is required โ the transitive buildTransitive shim copies onnxruntime-genai.dll automatically on Windows.
Add a runtime package only when you want a specific GPU provider:
# ๐ข NVIDIA GPU (CUDA):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.Cuda
# ๐ต Any Windows GPU โ AMD, Intel, NVIDIA (DirectML):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.DirectML
โ ๏ธ Add at most one GPU runtime package. Do not reference both
Microsoft.ML.OnnxRuntimeGenAI.CudaandMicrosoft.ML.OnnxRuntimeGenAI.DirectMLsimultaneously.If you use a GPU runtime package and want to disable the transitive CPU copy shim, set:
<ElBrunoLocalLLMsDisableCpuNativeCopy>true</ElBrunoLocalLLMsDisableCpuNativeCopy>in your application.csproj.
๐ The library defaults to
ExecutionProvider.Autoโ on Windows it tries DirectML โ CUDA โ CPU, and on Linux it tries CUDA โ CPU. No code changes needed.
Quick Start
using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;
// Create a local chat client (downloads Phi-3.5 mini on first run)
using var client = await LocalChatClient.CreateAsync();
var response = await client.GetResponseAsync([
new(ChatRole.User, "What is the capital of France?")
]);
Console.WriteLine(response.Text);
First Run
The first time you create a LocalChatClient, the model is downloaded from HuggingFace to your local cache directory (~2-4 GB). This typically takes 30-60 seconds depending on your internet connection.
Track download progress:
using var client = await LocalChatClient.CreateAsync(
new LocalLLMsOptions { Model = KnownModels.Phi35MiniInstruct },
progress: new Progress<ModelDownloadProgress>(p =>
{
var percent = (p.BytesDownloaded * 100) / p.TotalBytes;
Console.WriteLine($"{p.FileName}: {percent:F1}%");
})
);
Subsequent runs load instantly from cache (%LOCALAPPDATA%/ElBruno/LocalLLMs/models).
Skip auto-download if using a pre-downloaded model:
var options = new LocalLLMsOptions
{
Model = KnownModels.Phi35MiniInstruct,
ModelPath = "/path/to/local/model",
EnsureModelDownloaded = false
};
using var client = await LocalChatClient.CreateAsync(options);
Streaming
using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;
using var client = await LocalChatClient.CreateAsync(new LocalLLMsOptions
{
Model = KnownModels.Phi35MiniInstruct
});
await foreach (var update in client.GetStreamingResponseAsync([
new(ChatRole.System, "You are a helpful assistant."),
new(ChatRole.User, "Explain quantum computing in simple terms.")
]))
{
Console.Write(update.Text);
}
GPU Acceleration
By default, ExecutionProvider.Auto tries GPU first and falls back to CPU automatically:
// Use explicit GPU provider (fails if CUDA not installed; use Auto to fallback to CPU)
var options = new LocalLLMsOptions
{
ExecutionProvider = ExecutionProvider.Cuda
};
// Multi-GPU systems: select device ID
var options2 = new LocalLLMsOptions
{
ExecutionProvider = ExecutionProvider.Cuda,
GpuDeviceId = 1 // Use second GPU
};
Auto fallback behavior:
- Windows + DirectML available โ uses a Windows GPU through DirectML
- Windows + DirectML unavailable, CUDA available โ uses NVIDIA GPU
- Linux + CUDA available โ uses NVIDIA GPU
- GPU unavailable โ falls back to CPU (no errors, just slower)
โ ๏ธ CUDA note: ONNX Runtime GenAI
0.15.xexpects CUDA13.*, cuDNN9.*, and the latest Microsoft Visual C++ 2015-2022 runtime. When those native libraries are missing, provider diagnostics now surface the exact DLL mismatch or missing dependency instead of entering the failing native path.
See Troubleshooting: GPU Setup for debugging GPU issues.
Model Metadata
Inspect model capabilities at runtime โ context window size, model name, and vocabulary:
using var client = await LocalChatClient.CreateAsync();
var metadata = client.ModelInfo;
Console.WriteLine($"Model: {metadata?.ModelName}");
Console.WriteLine($"Context window: {metadata?.MaxSequenceLength}");
Console.WriteLine($"Vocab size: {metadata?.VocabSize}");
This is useful for prompt-length validation, adaptive chunking, and model selection logic.
Dependency Injection
builder.Services.AddLocalLLMs(options =>
{
options.Model = KnownModels.Phi35MiniInstruct;
options.ExecutionProvider = ExecutionProvider.DirectML;
});
// Inject IChatClient anywhere
public class MyService(IChatClient chatClient) { ... }
Error Handling
The library provides structured exception types for graceful error handling:
using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;
try
{
using var client = await LocalChatClient.CreateAsync();
var response = await client.GetResponseAsync([
new(ChatRole.User, "Your question here")
]);
}
catch (ExecutionProviderException ex)
{
// GPU/provider-specific error (no CUDA, DirectML not available, etc.)
Console.WriteLine($"Provider error: {ex.Message}");
}
catch (ModelCapacityExceededException ex)
{
// Prompt/response too long for model's context window
Console.WriteLine($"Capacity error: {ex.Message}");
// Solution: use a larger model or truncate the prompt
}
catch (InvalidOperationException ex)
{
// General operation error (model not found, download failed, etc.)
Console.WriteLine($"Operation error: {ex.Message}");
}
Observability
LocalChatClient emits generation lifecycle diagnostics through ActivitySource and Meter, both named ElBruno.LocalLLMs.
using ElBruno.LocalLLMs.Diagnostics;
builder.Services.AddOpenTelemetry()
.WithTracing(tracing => tracing.AddSource(LocalLLMsInstrumentation.ActivitySourceName))
.WithMetrics(metrics => metrics.AddMeter(LocalLLMsInstrumentation.MeterName));
By default, telemetry excludes prompt and completion text. Opt in only when you want content attached:
var options = new LocalLLMsOptions
{
CaptureTelemetryContent = true
};
See docs/observability.md for the lifecycle event contract, metric names, and Aspire wiring notes, and docs/cancellation.md for voice barge-in cancellation behavior.
Local Decisions
Sometimes you don't need prose โ you need a decision. ElBruno.LocalLLMs.Decisions runs a
System One model that returns a typed answer with its full probability distribution in a
single forward pass, fast enough to call on every request.
dotnet add package ElBruno.LocalLLMs.Decisions
The model runs in-process on ONNX Runtime โ no server to start and no Python. The weights are
downloaded from elbruno/laya-onnx on first use
(about 800 MB) and cached afterwards.
Ask several questions at once โ they share one forward pass, so four cost about what one costs:
using var client = new LayaOnnxDecisionClient();
DecisionResult result = await client.EvaluateAsync(
new DecisionRequest("My invoice charged me twice and I want my money back.")
.Choose("team", new Dictionary<string, string?>
{
["billing"] = "Payment, invoice and refund problems",
["technical"] = "Bugs, outages and API errors",
["sales"] = "Pricing, upgrades and new purchases"
})
.Score("urgency", new[] { "no rush", "normal", "urgent", "critical" })
.Ask("refund", "The customer is asking for a refund."));
ChoiceResult team = result.Choice("team");
// Gate on the probability so ambiguous tickets reach a human
// instead of being confidently misrouted.
string route = team.ChoiceOrNull(0.6) ?? "human-review";
Console.WriteLine(route); // billing
Console.WriteLine(result.Score("urgency").Score); // 1.67 of 3
Console.WriteLine(result.Probability("refund").IsTrue); // True
โ ๏ธ Laya's public checkpoints report uncalibrated confidence. Fit
DecisionThresholdagainst your own labelled examples before relying on a boolean verdict. Answers whose temperature bucket the checkpoint got wrong are flagged withCalibrationClampedโ see the Local Decisions Guide and the DecisionCalibration sample.
Cache Management
Inspect and manage the local model cache programmatically:
// Remove a model from the cache (no-op if not cached)
await LocalChatClient.DeleteModelFromCacheAsync(KnownModels.Phi35MiniInstruct);
// Or use a custom cache directory
await LocalChatClient.DeleteModelFromCacheAsync(
KnownModels.Phi35MiniInstruct,
cacheDirectory: @"D:\my-models");
// Get cached size in bytes for one model (0 if not downloaded)
long bytes = LocalChatClient.GetModelCacheSize(KnownModels.Phi35MiniInstruct);
Console.WriteLine($"Cached: {bytes / 1024 / 1024:N0} MB");
// List all cached models with size and last-modified date
var cached = LocalChatClient.ListCachedModels();
foreach (var repo in cached)
Console.WriteLine($"{repo.LocalDirectory} {repo.TotalSizeBytes / 1024 / 1024:N0} MB {repo.LastModified:yyyy-MM-dd}");
// Same APIs available on LocalVisionChatClient for vision models
await LocalVisionChatClient.DeleteModelFromCacheAsync(KnownModels.Fara15_9B);
long visionBytes = LocalVisionChatClient.GetModelCacheSize(KnownModels.Fara15_9B);
The default cache directory is %LOCALAPPDATA%/ElBruno/LocalLLMs/models (Windows) or ~/.local/share/ElBruno/LocalLLMs/models (Linux/macOS).
These operations delegate to ElBruno.HuggingFace.Downloader which provides the underlying DeleteCachedFilesAsync, GetCachedSize, and ListCachedRepos implementation.
Troubleshooting
GPU not working? Use ExecutionProvider.Cpu explicitly. See GPU Setup Validation.
Out of memory? Try a smaller model:
var options = new LocalLLMsOptions
{
Model = KnownModels.Qwen25_05BInstruct // 0.5B instead of 3.8B
};
Model download fails?
- Check your internet connection
- For private HuggingFace models, set the
HF_TOKENenvironment variable
For detailed troubleshooting, see docs/troubleshooting-guide.md.
Supported Models
| Tier | Model | Parameters | ONNX | ID |
|---|---|---|---|---|
| โช Tiny | TinyLlama-1.1B-Chat | 1.1B | โ Native | tinyllama-1.1b-chat |
| โช Tiny | SmolLM2-1.7B-Instruct | 1.7B | โ Native | smollm2-1.7b-instruct |
| โช Tiny | Qwen2.5-0.5B-Instruct | 0.5B | โ Native | qwen2.5-0.5b-instruct |
| โช Tiny | Qwen2.5-1.5B-Instruct | 1.5B | โ Native | qwen2.5-1.5b-instruct |
| โช Tiny | Gemma-2B-IT | 2B | โ Native | gemma-2b-it |
| โช Tiny | Gemma-4-E2B-IT | 5.1B (2B active) | ๐ Convert | gemma-4-e2b-it |
| โช Tiny | StableLM-2-1.6B-Chat | 1.6B | โ Native | stablelm-2-1.6b-chat |
| ๐ข Small | Phi-3.5 mini instruct | 3.8B | โ Native | phi-3.5-mini-instruct |
| ๐ข Small | Qwen2.5-3B-Instruct | 3B | โ Native | qwen2.5-3b-instruct |
| ๐ข Small | Llama-3.2-3B-Instruct | 3B | โ Native | llama-3.2-3b-instruct |
| ๐ข Small | Gemma-2-2B-IT | 2B | โ Native | gemma-2-2b-it |
| ๐ข Small | Gemma-4-E4B-IT | 8B (4B active) | ๐ Convert | gemma-4-e4b-it |
| ๐ก Medium | Qwen2.5-7B-Instruct | 7B | โ Native | qwen2.5-7b-instruct |
| ๐ก Medium | Qwen2.5-Coder-7B-Instruct | 7B | โ Native | qwen2.5-coder-7b-instruct |
| ๐ก Medium | Llama-3.1-8B-Instruct | 8B | โ Native | llama-3.1-8b-instruct |
| ๐ก Medium | Mistral-7B-Instruct-v0.3 | 7B | โ Native | mistral-7b-instruct-v0.3 |
| ๐ก Medium | Gemma-2-9B-IT | 9B | โ Native | gemma-2-9b-it |
| ๐ก Medium | Gemma-4-12B-IT | 12B | ๐ Convert | gemma-4-12b-it |
| ๐ก Medium | Phi-4 | 14B | โ Native | phi-4 |
| ๐ก Medium | DeepSeek-R1-Distill-Qwen-14B | 14B | โ Native | deepseek-r1-distill-qwen-14b |
| ๐ก Medium | Mistral-Small-24B-Instruct | 24B | โ Native | mistral-small-24b-instruct |
| ๐ด Large | Qwen2.5-14B-Instruct | 14B | โ Native | qwen2.5-14b-instruct |
| ๐ด Large | Qwen2.5-32B-Instruct | 32B | โ Native | qwen2.5-32b-instruct |
| ๐ด Large | Llama-3.3-70B-Instruct | 70B | โ ONNX | llama-3.3-70b-instruct |
| ๐ด Large | Mixtral-8x7B-Instruct-v0.1 | 8x7B | โ Native | mixtral-8x7b-instruct-v0.1 |
| ๐ด Large | DeepSeek-R1-Distill-Llama-70B | 70B | โ Native | deepseek-r1-distill-llama-70b |
| ๐ด Large | Command-R (35B) | 35B | โ Native | command-r-35b |
| ๐ด Large | Gemma-4-26B-A4B-IT | 25.2B (3.8B active) | ๐ Convert | gemma-4-26b-a4b-it |
| ๐ด Large | Gemma-4-31B-IT | 30.7B | ๐ Convert | gemma-4-31b-it |
| ๐ฃ Next-Gen | Qwen3-14B-Instruct | 14.77B | โ Native | qwen3-14b-instruct |
| ๐ง GPT-OSS | GPT-OSS 20B (CPU INT4) | 21B (3.6B active, MoE) | โ Native | gpt-oss-20b |
| ๐ง GPT-OSS | GPT-OSS 20B (CUDA INT4) | 21B (3.6B active, MoE) | โ Native | gpt-oss-20b-cuda |
| ๐ค Agentic | MagenticBrain | ~14.77B | โ Native | magentic-brain |
| ๐๏ธ VLM | Fara 1.5-9B | ~9.4B | โ Native | fara-1.5-9b |
๐ Convert = Use the conversion scripts in
scripts/to export ONNX locally before running the model.ยน MagenticBrain ONNX: Native ONNX hosted at
elbruno/MagenticBrain-onnx(INT4 quantized). Auto-downloads whenEnsureModelDownloaded=true.ยฒ Fara 1.5-9B ONNX:
elbruno/Fara1.5-9B-onnxnow includes the validated multimodal package (qwen3vl-vision.onnx,qwen3vl-embedding.onnx, patchedgenai_config.json, and ORT-compatibleprocessor_config.json). See ONNX Conversion โ Fara.ยณ GPT-OSS 20B: Apache-2.0, from the official
onnxruntime/gpt-oss-20b-onnxrepository. The CPU INT4 variant is a ~12 GB download, and because GPT-OSS is a mixture-of-experts model, CPU inference is slow โ prefergpt-oss-20b-cudawithMicrosoft.ML.OnnxRuntimeGenAI.Cudawhen a GPU is available. GPT-OSS reasons before answering; that chain-of-thought is filtered out and never surfaced, per the model card. Reasoning depth is controlled byLocalLLMsOptions.ReasoningEffort.
Fine-Tuned Models
Pre-trained variants optimized for specific tasks. A fine-tuned 0.5B model often matches or exceeds a base 1.5B on its specialized task.
| Model | Size | Task | HuggingFace ID |
|---|---|---|---|
| Qwen2.5-0.5B-ToolCalling | ~1 GB | Tool/function calling | elbruno/Qwen2.5-0.5B-LocalLLMs-ToolCalling |
| Qwen2.5-0.5B-RAG | ~1 GB | RAG with citations | elbruno/Qwen2.5-0.5B-LocalLLMs-RAG |
| Qwen2.5-0.5B-Instruct | ~1 GB | General-purpose | elbruno/Qwen2.5-0.5B-LocalLLMs-Instruct |
See the Supported Models Guide for detailed model cards, performance benchmarks, and selection guidance.
Samples
| Sample | Description |
|---|---|
| HelloChat | Minimal console chat |
| StreamingChat | Token-by-token streaming |
| MultiModelChat | Switch models at runtime |
| DependencyInjection | ASP.NET Core DI registration |
| ToolCallingAgent | Function calling and tool use |
| FineTunedToolCalling | Fine-tuned model for improved tool calling |
| RagChatbot | RAG pipeline with document retrieval |
| ZeroCloudRag | Zero-cloud RAG pipeline with real local embeddings and LLM inference |
| BitNetChat | BitNet 1.58-bit model chat completion |
| BitNetPerformance | Performance benchmark: BitNet vs ONNX models |
| MagenticBrainAgent | Multi-agent orchestration loop using Qwen3/MagenticBrain |
| FaraVisionAgent | Vision-language model (Fara 1.5-9B) image+text inference |
| GptOssChat | GPT-OSS 20B chat, streaming, reasoning effort, and tool calling |
| MagenticUIServer | ASP.NET Core + SignalR multi-agent server (FileSurfer, WebFetcher, Coder) |
| ConsoleAppDemo | Interactive console application |
| LocalDecisions | Support-ticket triage with a local System One decision model |
| DecisionCalibration | Calibration clamping and fp16 batch variance, made visible |
๐ Reference App: ElBruno.MagenticUI โ full Blazor Server port of microsoft/magentic-ui running locally with this library.
Requirements
- .NET 10.0
- CPU (default), NVIDIA GPU (CUDA), or Windows GPU (DirectML)
- ~2-8 GB disk space per model (depending on size and quantization)
Building from Source
git clone https://github.com/elbruno/ElBruno.LocalLLMs.git
cd ElBruno.LocalLLMs
dotnet restore ElBruno.LocalLLMs.slnx
dotnet build ElBruno.LocalLLMs.slnx
dotnet test ElBruno.LocalLLMs.slnx
Run integration tests (downloads real models โ requires internet):
RUN_INTEGRATION_TESTS=true dotnet test ElBruno.LocalLLMs.slnx
Integration tests validate the full lifecycle (download โ infer โ cache hit โ delete) for all 35 supported models. See docs/tests/README.md for details.
Documentation
- Getting Started โ installation, first steps, configuration
- Supported Models โ full model reference with tiers, specs, decision tree
- BitNet Guide โ setup and usage of 1.58-bit BitNet models
- Cancellation Guide โ streaming cancellation semantics and voice barge-in usage
- Observability โ OpenTelemetry, lifecycle events, metrics, Aspire wiring
- Blazor Components Guide โ the 7 Razor components and how to wire them
- Local Decisions Guide โ typed decision models, calibration caveats, thresholds
- Architecture โ design decisions and internal structure
- Samples Guide โ walkthrough of each sample application
- Benchmarks โ how to run and interpret performance benchmarks
- Fine-Tuning Guide โ using and training fine-tuned models
- ONNX Conversion โ converting HuggingFace models to ONNX format
- ONNX Conversion โ Fara VLM โ converting Fara 1.5-9B vision-language model
- Blog: Fara End-to-End Support โ TL;DR, motivation, and C# usage for Fara 1.5-9B
- Blog Companion: Fara Assets โ suggested titles/tags and image prompts
- Publishing โ NuGet package publishing with OIDC
- Contributing โ how to contribute
- Integration Test Reports โ how to run E2E lifecycle tests and interpret per-run markdown reports
- Changelog โ version history
๐ค Contributing
Contributions are welcome! Please:
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
๐ License
This project is licensed under the MIT License โ see the LICENSE file for details.
๐ About the Author
Hi! I'm ElBruno ๐งก, a passionate developer and content creator exploring AI, .NET, and modern development practices.
Made with โค๏ธ by ElBruno
If you like this project, consider following my work across platforms:
- ๐ป Podcast: No Tienen Nombre โ Spanish-language episodes on AI, development, and tech culture
- ๐ป Blog: ElBruno.com โ Deep dives on embeddings, RAG, .NET, and local AI
- ๐บ YouTube: youtube.com/elbruno โ Demos, tutorials, and live coding
- ๐ LinkedIn: @elbruno โ Professional updates and insights
- ๐ Twitter: @elbruno โ Quick tips, releases, and tech news
๐ Acknowledgments
- ONNX Runtime GenAI โ inference engine
- BitNet / bitnet.cpp โ 1.58-bit ternary model inference
- Microsoft.Extensions.AI โ IChatClient interface
- Hugging Face โ model hosting and community
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- ElBruno.LocalLLMs (>= 0.22.0)
- Microsoft.Data.Sqlite (>= 9.0.3)
- Microsoft.Extensions.AI.Abstractions (>= 10.8.3)
- SQLitePCLRaw.bundle_e_sqlite3 (>= 2.1.12)
NuGet packages (1)
Showing the top 1 NuGet packages that depend on ElBruno.LocalLLMs.Rag:
| Package | Downloads |
|---|---|
|
ElBruno.LocalLLMs.BlazorComponents
Blazor UI components for ElBruno.LocalLLMs โ model status cards, chat boxes, download management, environment diagnostics, RAG playground and more. |
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.22.0 | 7 | 9/30/2026 |
| 0.21.0 | 169 | 8/11/2026 |
| 0.20.12 | 111 | 8/11/2026 |
| 0.20.11 | 122 | 8/11/2026 |
| 0.20.10 | 128 | 8/10/2026 |
| 0.20.9 | 135 | 8/2/2026 |
| 0.20.8 | 149 | 8/1/2026 |
| 0.20.7 | 121 | 8/1/2026 |
| 0.20.6 | 127 | 8/1/2026 |
| 0.20.5 | 113 | 7/30/2026 |
| 0.20.4 | 115 | 7/29/2026 |
| 0.20.3 | 114 | 7/28/2026 |
| 0.20.1 | 113 | 7/24/2026 |
| 0.20.0 | 109 | 7/23/2026 |
| 0.19.1 | 134 | 7/3/2026 |
| 0.19.0 | 131 | 7/3/2026 |
| 0.18.0 | 136 | 6/5/2026 |
| 0.17.0 | 119 | 6/3/2026 |
| 0.16.0 | 127 | 4/17/2026 |
| 0.15.0 | 119 | 4/16/2026 |