Model-Agnostic AI Architecture: Swap LLMs Safely
Model-agnostic AI architecture means the LLM is a replaceable component behind one adapter. You keep versioned I/O contracts, a tools/MCP allowlist with policy outside the prompt, retrieval filtered by ACL before the model sees text, a golden-set eval harness, and traces that record model, prompt-pack, tools, tokens, cost, and latency. You change Azure OpenAI deployment name, prompt-pack version, or SKU only after golden-set, regression, cost/latency, safety checks, and canary gates pass.
Short picture
The body is what you own (contracts, tools, ACL, evals, logs). The head is the model behind one adapter. Swap the head only after offline gates pass. The rest of this article is that socket for Support Answer on Azure.
Use case: Support Answer on Azure
A support agent asks: “Why was my order late?”
- Support UI sends
tenantId,userId,orderId, question, conversation id. - API / BFF (ASP.NET Core) authenticates, rate-limits, posts
SupportAnswerRequestto the AI microservice. - The microservice queries Azure AI Search with an ACL filter, builds a prompt pack, calls Azure OpenAI through a model adapter.
- The model may propose tools (e.g.
get_shipment_events). The tool/MCP gateway allowlists, applies policy, audits, then executes — not the prompt text. - The response is validated against a versioned schema and returned to the BFF.
- App Insights records deployment, prompt-pack version, tool ids, tokens, estimated cost, latency.
Auth (Azure): Prefer managed identity on App Service / ACA with DefaultAzureCredential to call Azure OpenAI (and AI Search when RBAC is enabled). Key Vault holds secrets only when you still need them (legacy API keys, third-party tool credentials)—accessed by the host identity, not by the model adapter as a special path. Do not put keys in prompts or app settings in clear text.
Swap means a new Azure OpenAI deployment and/or prompt-pack version. Callers still use IModelAdapter. Cutover waits on the offline gate pipeline.
1. Versioned I/O contract
Do not trust free-form prose from the model as the service boundary. Define a contract, version it, reject invalid payloads.
Serialize with camelCase JSON if the wire format is JSON (PropertyNamingPolicy = JsonNamingPolicy.CamelCase); keep CLR names PascalCase in C#.
C# records (v1)
public sealed record SupportAnswerRequest(
string SchemaVersion, // "support-answer/1"
string TenantId,
string UserId,
string OrderId,
string Question,
string? ConversationId);
public sealed record SupportAnswerResponse(
string SchemaVersion, // "support-answer/1"
string Intent, // e.g. order_delay_explain
string Answer,
IReadOnlyList<Citation> Citations,
IReadOnlyList<ToolCallProposal> ToolCalls, // model proposals only
bool Refuse,
string? RefuseReason);
public sealed record Citation(string DocumentId, string Snippet, double? Score);
// Model returned this — arguments only. No execution result here.
public sealed record ToolCallProposal(string Name, string ArgumentsJson);
// Gateway wrote this after allowlist + policy + execute.
public sealed record ToolExecutionResult(
string Name,
string ArgumentsJson,
string? ResultJson,
bool Allowed,
string? DenyReason);
JSON Schema sketch (response)
{
"$id": "support-answer/1/response",
"type": "object",
"required": ["schemaVersion", "intent", "answer", "citations", "toolCalls", "refuse"],
"properties": {
"schemaVersion": { "const": "support-answer/1" },
"intent": { "type": "string", "minLength": 1 },
"answer": { "type": "string" },
"citations": {
"type": "array",
"items": {
"type": "object",
"required": ["documentId", "snippet"],
"properties": {
"documentId": { "type": "string" },
"snippet": { "type": "string" },
"score": { "type": "number" }
}
}
},
"toolCalls": {
"type": "array",
"items": {
"type": "object",
"required": ["name", "argumentsJson"],
"properties": {
"name": { "type": "string" },
"argumentsJson": { "type": "string" }
}
}
},
"refuse": { "type": "boolean" },
"refuseReason": { "type": ["string", "null"] }
}
}
The model fills the form. Your code validates and maps to HTTP. Execution results live in the gateway audit trail / traces, not inside the model’s proposed toolCalls array unless you deliberately add a separate response field after the tool loop.
2. Model adapter (swap without touching callers)
Feature code depends on an interface, not on Azure.AI.OpenAI types in every handler.
// Sketch — production adapter wraps Azure OpenAI chat completions.
public interface IModelAdapter
{
Task<ModelCompletion> CompleteAsync(ModelRequest request, CancellationToken ct);
}
public sealed record ModelRequest(
string PromptPackVersion,
IReadOnlyList<ChatMessage> Messages,
IReadOnlyList<ToolDefinition> Tools,
ResponseFormat ResponseFormat);
public sealed record ModelCompletion(
string RawText,
IReadOnlyList<ToolCallProposal> ToolCalls,
ModelUsage Usage,
string DeploymentName,
string PromptPackVersion);
What swaps in config (not in callers):
| Knob | Where it lives | Example |
|---|---|---|
| Deployment name | AzureOpenAi:DeploymentName | support-gpt-4o → support-gpt-4.1-mini |
| API version | options | 2024-10-21 |
| SKU / capacity | Azure portal + cost label in options | Standard → provisioned |
| Prompt pack | blob/app config version id | support-answer.promptpack@3 |
Register options or adapter implementation in DI. No ChatClient construction inside the Support Answer handler.
3. Tools / MCP: allowlist and policy outside the prompt
The model may propose get_shipment_events. It is not the authority on whether that tool runs.
public sealed class ToolGateway
{
private readonly HashSet<string> _allowlist = new(StringComparer.Ordinal)
{
"get_shipment_events",
"get_order_status"
};
public async Task<ToolExecutionResult> ExecuteAsync(
ToolCallProposal proposal,
CallerContext caller,
CancellationToken ct)
{
if (!_allowlist.Contains(proposal.Name))
return new ToolExecutionResult(proposal.Name, proposal.ArgumentsJson, null, false, "tool_not_allowlisted");
if (!await _policy.CanInvokeAsync(proposal.Name, caller, ct))
return new ToolExecutionResult(proposal.Name, proposal.ArgumentsJson, null, false, "policy_denied");
await _audit.WriteAsync(caller, proposal, ct);
var resultJson = await _executors[proposal.Name].RunAsync(proposal, caller, ct);
return new ToolExecutionResult(proposal.Name, proposal.ArgumentsJson, resultJson, true, null);
}
}
Prompt text can prefer citing shipment events. It must not be the only place that forbids refunds. Keep refunds off the allowlist until product and policy say otherwise. Same gateway pattern for MCP tools.
Related: Evaluating AI Agents in Production.
4. Retrieval + ACL before the prompt
Filter by tenant and user rights before chunks enter the prompt. The model is not your ACL layer.
A Search $filter is not Entra authorization. Callers must already be authenticated; the filter only narrows indexed documents for that principal’s attributes you trust from the token or your user store.
// Escape OData string literals (quotes). Prefer typed builders if you have them.
static string ODataEscape(string value) => value.Replace("'", "''");
var filter =
$"tenantId eq '{ODataEscape(tenantId)}' and " +
$"(visibility eq 'public' or sharedWith/any(u: u eq '{ODataEscape(userId)}'))";
var options = new SearchOptions
{
Filter = filter,
Size = 8,
IncludeTotalCount = false
};
options.Select.Add("documentId");
options.Select.Add("content");
options.Select.Add("orderId");
var results = await _searchClient.SearchAsync<SupportChunk>(question, options, ct);
// Only these chunks enter the prompt pack.
If a document fails the filter, it never appears in the prompt. Do not “ask the model to ignore confidential rows.”
5. Eval harness
A golden set is a fixed list of real support questions with expected outcomes — not “sounds helpful.”
Case shape
{
"id": "delay-carrier-scan-missing",
"request": {
"schemaVersion": "support-answer/1",
"tenantId": "t-demo",
"userId": "u-agent-1",
"orderId": "ORD-10042",
"question": "Why was my order late?"
},
"expect": {
"intent": "order_delay_explain",
"mustCiteDocumentIds": ["shp-10042-events"],
"forbidTools": ["create_refund"],
"refuse": false,
"answerMustInclude": ["carrier", "scan"]
}
}
Pass / fail (examples)
| Check | Pass if |
|---|---|
| Schema | Response validates against support-answer/1 |
| Intent | Exact or approved alias map |
| Citations | Required doc ids present; no doc outside ACL fixture |
| Tools | No forbidden tool proposals executed by the gateway |
| Refuse | Matches expect when the case should refuse |
| Regression | Success rate ≥ current deployment − tolerance (e.g. 2 pp) |
Run the same harness against current and candidate deployments. Attach the report to the change PR. Broader failure-mode notes: AI Architecture That Holds in Production.
6. Observability fields
Every Support Answer completion should emit (App Insights custom dimensions / spans):
| Field | Why |
|---|---|
model.deployment | Which Azure OpenAI deployment answered |
model.apiVersion | API contract used |
promptPack.version | Which instructions/examples |
schema.version | support-answer/1 |
tools.proposed / tools.executed / tools.denied | Drift and policy denials |
retrieval.documentIds | Citation audit |
usage.inputTokens / usage.outputTokens | Cost |
cost.estimatedUsd | Budget gate |
latency.ms | p95 budget |
canary.bucket | Which traffic slice |
traceId | Join to BFF and tool audit |
Without these, “the new model feels worse” is not actionable.
What you may swap
| Piece | What it is | Why it moves |
|---|---|---|
| Model deployment | Azure OpenAI deployment name | Quality, price, context limits |
| Prompt pack | Versioned instructions + few-shots | Behaviour tune without code |
| SKU / tier | Capacity and $/token | Budget and throughput |
All three go through the adapter and config. Feature handlers stay on IModelAdapter + SupportAnswerRequest/Response.
Validation gates (offline change pipeline)
Gates are not on the live request path. Input to the pipeline is a candidate deployment + prompt-pack version. Fail means keep the current defaults.
- Golden set — candidate vs pass/fail table above.
- Regression — compare candidate to current on the same set.
- Cost / p95 — estimated $/successful answer and p95 latency within envelope.
- Safety check — re-run fixtures that prove the gateway still denies forbidden tools and ACL filters still exclude private docs. This confirms your controls; it is not a claim that the model is “safe” in the abstract.
- Canary — small live traffic, then cutover or hold.
Canary (accurate for Azure)
Azure OpenAI deployments do not offer a built-in traffic weight between two deployments for your app’s requests. You create a candidate deployment, then split traffic in the application (feature flag / sticky canary.bucket on tenantId) or at API Management. Kill switch = force config back to the current deployment name (seconds), not a full business-logic redeploy.
App Service deployment slots swap app revisions. Useful when adapter/options ship with the app. They do not replace golden-set proof for a new model. Prefer: prove on golden set → canary via app/APIM → promote default config.
What NOT to glue into prompts
- ACL rules as the only control
- Tool permissions without allowlist + policy
- Raw secrets or API keys
- Unversioned portal-only system prompts with no prompt-pack id in logs
- Provider SDK calls copy-pasted into every feature
- Success criteria that only say “be helpful”
Those belong in contracts, gateway, search filters, config, and evals.
When to change models
Change when all are true:
- Measured need (cost, quality, or hard limit on the current deployment).
- Candidate beats current on the golden set for the metrics that matter, or matches quality at clearly lower cost.
- Cost and p95 budgets still hold.
- Canary stays boring for the agreed window (no spike in bad tool calls or empty citations).
Do not change because a launch blog named a new model.
PR checklist
- [ ] All model calls go through
IModelAdapter - [ ] Responses validate against versioned schema (
JsonNamingPolicy.CamelCaseif JSON wire) - [ ] Tools/MCP allowlist unchanged by the swap (or policy PR linked)
- [ ] Retrieval ACL filter runs before prompt construction (escaped; not a substitute for Entra auth)
- [ ] Golden-set + regression report attached
- [ ] p95 latency and $/success within budget
- [ ] Canary plan + kill switch named (app flag or APIM — not AOAI weights)
- [ ] Prompt-pack version and deployment name recorded in traces
FAQ
What is model-agnostic AI architecture?
A design where the LLM is swappable. Stable parts: I/O contracts, tool allowlists and policy, retrieval ACL, eval harness, and traces. Changing parts: model deployment, prompt pack, price tier.
What should stay stable when you change models?
Contracts, tools/MCP policy and audit, retrieval + ACL, golden-set evals, and logs that record deployment, prompt pack, retrieved ids, tools, latency, and cost.
What checks do you run before swapping an LLM?
(1) Golden set (2) regression vs current deployment (3) cost and p95 latency (4) safety fixtures for gateway + ACL (5) canary with a kill switch, then cutover.
When should you change models?
Only with a measured need, a win or equal quality at lower cost on the golden set, budgets intact, and a boring canary window.
Is model-agnostic the same as multi-model routing?
No. Routing picks a model per request. Model-agnostic means any chosen model still plugs into the same contracts and gates. You can do both; start with the socket.
Closing
Prove the swap on the offline gate pipeline, then promote the head. Support Answer on Azure is one concrete path.
Related notes: case studies. Architecture questions: contact.
— AI Solution Architect · Sydney