.NET + AI
ASP.NET Core vs Python FastAPI for AI APIs: Which One Should Serve Your Model?
ASP.NET Core vs Python FastAPI for AI APIs — the same streaming LLM endpoint built in both, and a decision framework based on where your model actually runs.
ASP.NET Core vs Python FastAPI for AI APIs
Here is the same AI endpoint written twice. It takes a prompt, sends it to a local model running in Ollama, and streams tokens back to the client as Server-Sent Events.
FastAPI:
import json
import httpx
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from pydantic import BaseModel, Field
OLLAMA_URL = "http://localhost:11434/api/chat"
app = FastAPI()
client = httpx.AsyncClient(timeout=120.0)
class ChatRequest(BaseModel):
message: str = Field(min_length=1, max_length=4000)
model: str = "llama3.1"
@app.post("/chat")
async def chat(req: ChatRequest):
async def tokens():
payload = {
"model": req.model,
"messages": [{"role": "user", "content": req.message}],
"stream": True,
}
async with client.stream("POST", OLLAMA_URL, json=payload) as resp:
resp.raise_for_status()
async for line in resp.aiter_lines():
if not line:
continue
chunk = json.loads(line)
if chunk.get("done"):
break
yield f"data: {json.dumps(chunk['message']['content'])}\n\n"
return StreamingResponse(tokens(), media_type="text/event-stream")ASP.NET Core (.NET 10, Minimal API):
using System.ComponentModel.DataAnnotations;
using System.Text.Json;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddValidation();
builder.Services.AddHttpClient("ollama", c =>
{
c.BaseAddress = new Uri("http://localhost:11434");
c.Timeout = TimeSpan.FromMinutes(2);
});
var app = builder.Build();
app.MapPost("/chat", async (ChatRequest req, IHttpClientFactory factory,
HttpContext ctx, CancellationToken ct) =>
{
ctx.Response.ContentType = "text/event-stream";
var http = factory.CreateClient("ollama");
var payload = new
{
model = req.Model,
messages = new[] { new { role = "user", content = req.Message } },
stream = true
};
using var request = new HttpRequestMessage(HttpMethod.Post, "/api/chat")
{
Content = JsonContent.Create(payload)
};
using var response = await http.SendAsync(
request, HttpCompletionOption.ResponseHeadersRead, ct);
response.EnsureSuccessStatusCode();
using var reader = new StreamReader(await response.Content.ReadAsStreamAsync(ct));
while (await reader.ReadLineAsync(ct) is { } line)
{
if (line.Length == 0) continue;
using var doc = JsonDocument.Parse(line);
if (doc.RootElement.GetProperty("done").GetBoolean()) break;
var token = doc.RootElement.GetProperty("message").GetProperty("content").GetString();
await ctx.Response.WriteAsync($"data: {JsonSerializer.Serialize(token)}\n\n", ct);
await ctx.Response.Body.FlushAsync(ct);
}
});
app.Run();
record ChatRequest(
[Required, StringLength(4000, MinimumLength = 1)] string Message,
string Model = "llama3.1");Both work. Both validate input, stream tokens, and hold the connection open without blocking a thread while the model thinks. The FastAPI version is shorter; the C# version is more explicit about timeouts, cancellation, and HTTP client lifetime.
That's the honest summary of the ASP.NET Core vs Python FastAPI for AI APIs debate at the endpoint level: for a thin API in front of a model, the framework barely matters. The model takes seconds to respond. Framework overhead is measured in microseconds. Neither framework is your bottleneck.
The real decision is somewhere else — and it's mostly about where the model runs.
The question that actually decides it
Ask this first: does the model execute inside your API process, or behind an HTTP call?
Case A — model behind an API Case B — model in-process
Client Client
↓ ↓
Your API ──HTTP──▶ OpenAI / Claude / Your API
│ Ollama / vLLM │
├─ auth, quotas ├─ PyTorch / transformers
├─ RAG retrieval ├─ tokenizer, GPU memory
├─ SQL, business rules └─ model weights loaded at startup
└─ logging, billingCase A is most production AI today. Your API is an orchestrator: it authenticates the user, pulls context from a database or vector store, builds a prompt, calls a hosted or self-hosted model server, validates the output, and logs everything. That's a normal backend job with an unusually slow dependency. Either framework handles it well, so pick on team, ecosystem, and the systems you already run.
Case B is when you load weights yourself — a fine-tuned classifier, a custom embedding
model, a reranker, anything built on PyTorch or Hugging Face transformers. Here Python wins
outright. The training and inference ecosystem lives there, and fighting that is a waste of
engineering time.
Everything below is detail on top of that split.
Where FastAPI is the better choice
The ML ecosystem is Python-first. PyTorch, transformers, sentence-transformers,
LangChain, LlamaIndex, most evaluation tooling, and nearly every research repo you'll want to
borrow from ship Python first. If your roadmap includes "we'll fine-tune this later" or "let's
try that new reranker from last week's paper," FastAPI keeps you one pip install away.
Pydantic doubles as your LLM output schema. The same BaseModel that validates the
request body can define the structured JSON you want back from the model:
from pydantic import BaseModel
class TicketTriage(BaseModel):
category: str
priority: int
summary: str
# Use TicketTriage as the response schema for a structured-output call,
# then validate the model's reply with TicketTriage.model_validate_json(raw)One type definition covers the API contract, the OpenAPI docs FastAPI generates, and the model's output validation. That's a real productivity gain for prompt-heavy services.
Prototype speed. A data scientist can turn a notebook into a deployed endpoint without learning a second language. For internal tools and early-stage experiments, that matters more than anything in the next section.
Where ASP.NET Core is the better choice
CPU work around the model. AI APIs do more than wait. They chunk documents, parse PDFs,
tokenize, compute similarity scores, and serialize large payloads. In Python, CPU-bound code
inside an async handler blocks the event loop for every request on that worker. The standard
fix is more worker processes or offloading to a thread/process pool — which works, but each
process has its own memory, and anything loaded in-process gets loaded once per worker.
ASP.NET Core runs on a multi-threaded thread pool in a single process. CPU work spreads across cores without extra orchestration, and shared state (caches, connection pools, loaded configuration) exists once.
Long-lived streaming connections are well-trodden ground. Notice what the C# endpoint got
for free: the CancellationToken is tied to the HTTP request, so if the user closes the tab,
the upstream call to the model is cancelled too. You stop paying for tokens nobody will read.
IHttpClientFactory handles connection pooling and DNS refresh. These are defaults, not
add-ons.
It sits next to the rest of your business. Most AI features aren't standalone — they read from SQL Server, respect existing auth, write audit logs, and enforce the same permissions as the rest of the app. If that app is already .NET, putting the AI endpoint in the same stack means one deployment pipeline, one auth model, one set of logging and health checks. Our AI Database Agent is built on ASP.NET Core for exactly this reason: the hard part is SQL Server schema access, query validation, and safe execution, not the model call.
The .NET AI libraries caught up. Microsoft.Extensions.AI gives you an IChatClient
abstraction that works across providers, Semantic Kernel covers orchestration and tool
calling, and OpenAI ships an official .NET SDK. For Case A workloads, you are no longer
reaching for a Python-only library to talk to a model. See
Build RAG with .NET and
Run Ollama with .NET for working examples.
Deployment is a single artifact. dotnet publish produces a self-contained output you can
drop on a VPS behind Nginx and run as a systemd service. No virtualenv, no system Python
version drift, no dependency resolver surprises at deploy time. Python AI images that pull in
PyTorch and CUDA libraries get large quickly; a .NET API that calls a model over HTTP stays
small.
Side by side
| Concern | FastAPI | ASP.NET Core |
|---|---|---|
| Loading model weights in-process | Native (PyTorch, transformers) | Possible via ONNX Runtime, but a narrower ecosystem |
| Calling hosted / self-hosted models | Excellent | Excellent |
| Request validation | Pydantic, in the handler signature | DataAnnotations + AddValidation() (.NET 10) |
| Structured LLM output | Pydantic models reused as schemas | Records + System.Text.Json |
| CPU-bound work per request | Blocks the event loop unless offloaded | Thread pool across cores |
| Scaling on one box | Multiple worker processes | One process, multi-threaded |
| Streaming responses | StreamingResponse + async generator | Write + flush, or return IAsyncEnumerable<T> |
| Client-disconnect cancellation | Needs explicit handling | CancellationToken bound to the request |
| OpenAPI docs | Built in, zero config | Built-in OpenAPI document generation |
| Deploy artifact | venv or container image | Self-contained publish output |
| Best fit | ML teams, in-process models, research | Enterprise backends, orchestration, existing .NET shops |
The pattern most teams end up with
When you need both — custom models and a serious backend — don't force one framework to do the other's job. Split them:
Client
↓
ASP.NET Core API ← auth, rate limits, RAG, SQL, business rules, billing
↓ HTTP (internal)
FastAPI inference service ← loads the model, exposes /embed, /classify, /rerank
↓
GPU / model weightsThe ASP.NET Core layer owns everything a user touches. The FastAPI service is a small, private inference worker with two or three endpoints and no business logic. Each side scales on its own terms: the API on CPU and connections, the inference service on GPU memory.
This is the same shape as calling OpenAI or Ollama — the model is just behind your own HTTP endpoint instead of someone else's. Which means your API code doesn't change when you swap a self-hosted model for a hosted one, or the reverse.
How to decide
Run through these in order and stop at the first one that applies:
- You load PyTorch or Hugging Face models in the request path → FastAPI for that service. Put it behind your main API if you have one.
- Your team is Python-only and the service is mostly prompts and glue → FastAPI. Rewriting skills is more expensive than any runtime difference.
- Your AI feature lives inside an existing .NET system (SQL Server, existing auth, existing deploy pipeline) → ASP.NET Core. Don't add a second stack for an HTTP call.
- The endpoint does meaningful CPU work per request (chunking, parsing, scoring) and has to hold many concurrent streams → ASP.NET Core.
- You need both custom models and a production backend → ASP.NET Core API in front, FastAPI inference service behind it.
Takeaways
- For an API that calls a model over HTTP, framework speed is noise next to model latency. Choose on ecosystem, team, and existing systems.
- If model weights load inside your process, use Python. Don't fight the ML ecosystem.
- FastAPI's edge is Pydantic and the Python ML libraries; ASP.NET Core's edge is multi-core concurrency, request-bound cancellation, and fitting into enterprise backends.
- When you need both, split them: ASP.NET Core as the product API, FastAPI as a private inference service.
- Whichever you pick, get the fundamentals right first — timeouts, streaming, cancellation, input validation, and output validation. Those decide whether your AI API survives production, not the language it's written in.
Continue reading
Related articles
How to Build Text-to-SQL with ASP.NET Core
A working Text-to-SQL pipeline in ASP.NET Core and SQL Server: a least-privilege database identity, schema-aware prompts, Microsoft.Extensions.AI with Ollama, a ScriptDom validator, a SHOWPLAN cost check and a repair loop — and the benchmark numbers each step bought.
Read article →AI Database Agent, Part 1: Naïve Text-to-SQL and Why It Fails Silently
The baseline every Text-to-SQL demo starts from — table names in, SQL out — measured against 32 real questions. 53% correct, seven silently wrong answers, and an UPDATE and a DROP TABLE sent to the database.
Read article →Run Ollama Locally and Connect It to .NET
Run a local LLM with Ollama and connect it to an ASP.NET Core application.
Read article →