Dev
The Complete Picture: Azure Container Apps + Dapr Reference Architecture for .NET
Dev.toUnited States · NORTH AMERICA
What does the finished stack buy for its floor, and which seam between its layers breaks first? Priced in West Europe, the layers in front of the code cost about $1,730 a month before the first contai...
What does the finished stack buy for its floor, and which seam between its layers breaks first? Priced in West Europe, the layers in front of the code cost about $1,730 a month before the first container starts, while the .NET API behind them costs $20 to $140 across 1M to 100M requests. So the compute bill cannot be the argument: on compute alone, Container Apps never beats Functions. The seams are where the series' own parts handed off: Part 4 maps the health endpoints only in Development while Part 7 probes them in production, and the daprAIConnectionString Part 5 told you to set splits the trace pipeline.
Seven parts built the stack one layer at a time. Stacked, it is Front Door Premium at the edge, API Management Standard v2 behind a Private Link origin, and an internal Container Apps environment in your own VNet running two ASP.NET Core services on .NET 10, each with a managed Dapr sidecar. Cosmos DB and a Service Bus Premium namespace sit behind private endpoints and take no keys, only managed identities. Aspire runs the two services and their sidecars on your laptop with one F5, and Terraform writes down everything else.
Every C# and HCL block below is lifted from the companion repo, trimmed but not edited. The Bicep fragment for the telemetry settings, the health endpoint fix, the dead-letter fix, the alert command and the KQL are not in the repo, and each one says so where it appears. One naming note: this article calls the services orders-api and inventory-api, the state store statestore and the broker orders-pubsub. The Aspire sample from Part 4 says order-service and inventory-service, and Part 3's sample calls its state store orderstore. The strings differ; the wiring is the same.
One request, seven hops
Follow one POST /orders from a partner system to Cosmos DB and back. The numbers on the diagram are the order it happens in:
1. Front Door ends TLS and runs the WAF. Premium, not Standard, for two reasons that each settle it alone: the managed rule sets are Premium only, and so are Private Link origins (tier comparison). The request leaves Front Door carrying an X-Azure-FDID header and nothing that proves who sent it. Part 6 is the long version of every hop up to the container.
2. Private Link into API Management. Standard v2 is the cheapest tier that takes an inbound private endpoint and also integrates outbound into a VNet, and this topology needs both. The gateway trusts the hop because of the network path and a check-header on X-Azure-FDID, not because of a credential.
3. API Management calls the app with its own identity. It validates the caller's token and applies per-caller limits, then authentication-managed-identity fetches a token for orders-api with the gateway's system-assigned identity. The environment's internal ingress routes on the host header, and Container Apps authentication checks the app role. Set "assignment required" on the service principal so identities without the role are refused a token. Part 6 lists that as something to verify on first deploy: Entra documents it for applications, never for managed identities.
4. orders-api asks inventory-api whether it can fill the order, through its sidecar. This is the handler from Part 3, trimmed:
private static async Task<IResult> CreateOrderAsync(
CreateOrderRequest request,
InventoryClient inventory,
DaprClient dapr,
ILogger<Order> logger,
CancellationToken cancellationToken)
{
var stockRequest = new StockCheckRequest(
request.OrderId,
[.. request.Lines.Select(line => new StockLine(line.Sku, line.Quantity))]);
// Service invocation. The failure is a value, so both interesting branches (inventory
// said no; inventory could not be asked) are visible in the same switch.
var stock = await inventory.CheckStockAsync(stockRequest, cancellationToken);
switch (stock)
{
case Result<StockCheckResponse>.Failure(var error):
return ToProblem(error);
case Result<StockCheckResponse>.Success({ Available: false } answer):
return Results.Conflict(new OutOfStockResponse(request.OrderId, answer.Shortfalls));
}
var order = new Order(
request.OrderId,
request.CustomerId,
request.Lines,
request.Lines.Sum(line => line.UnitPrice * line.Quantity),
OrderStatus.Placed,
DateTimeOffset.UtcNow,
ConfirmedAt: null);
await dapr.SaveStateAsync(
StateStore.Name,
order.OrderId,
order,
metadata: OrderPartition.For(order.CustomerId),
cancellationToken: cancellationToken);
return Results.Created($"/orders/{order.CustomerId}/{order.OrderId}", order);
}
InventoryClient wraps an HttpClient whose base address is the Dapr app id, not a host. The two sidecars resolve each other inside the environment and wrap the call in mutual TLS without any configuration from you, and the platform injects APP_API_TOKEN so inventory-api can check that the call came through its own sidecar. Nothing checks which app sent it: Dapr's access control policies live in the Configuration spec, which Container Apps does not support, so any Dapr-enabled app in the environment can call inventory-api.
5. SaveStateAsync writes to Cosmos DB. The code names a component, StateStore.Name, and never an account. The component carries the account URL plus azureClientId, the client ID of orders-api's user-assigned identity, and the grant behind it is the data-plane role Cosmos DB Built-in Data Contributor. That role does not come from azurerm_role_assignment; Part 7 covers the resource that actually grants it, and the green apply you get from the one that does not.
6. The order is announced on a Service Bus topic. Part 3's handler stops at the write; the publish is one line in Part 2's sample, after the same save:
await dapr.PublishEventAsync(PubSubName, OrderCreatedTopic, orderCreated, cancellationToken);
The orders-pubsub component is pubsub.azure.servicebus.topics, authenticated with the same azureClientId pattern. The publisher's identity holds Azure Service Bus Data Sender, the subscriber's holds Data Receiver, and consumerID defaults to the app id and becomes the subscription name, which Part 5 showed matters as soon as two runtimes share a topic. On the receiving side, inventory-api's sidecar pulls the message and POSTs it to the subscription route, while a KEDA azure-servicebus rule scales inventory-api on subscription depth. The scale rule needs an identity reference of its own, even when it names the same identity the sidecar uses, or it cannot read the depth it scales on.
7. Every span goes to the environment's OpenTelemetry agent. The app's spans and the daprd spans leave each replica over OTLP to one managed agent per environment, which forwards them to Application Insights. It is the one hop in the chain that does not run on managed identity: the agent authenticates with a connection string, so the Application Insights resource has to keep local authentication switched on. The operations section below is mostly about this hop.
Seven hops, and the identity at each one:
The image pull does not appear on the diagram and uses the same user-assigned identity with AcrPull on the registry.
Who owns what: Aspire, Dapr, Container Apps, Terraform
Four tools touch this stack, and each one owns a different stage of it.
Aspire owns dev time. The AppHost starts both services, both sidecars and a Redis-backed state store with one F5, and hands every process its OTLP endpoint so the dashboard draws one trace across all of them. It stops at the deployment boundary: nothing in Azure runs the AppHost, as Part 4 made clear. What survives into production from Aspire is ServiceDefaults, compiled into every service.
Dapr owns the contract between the app and the platform. Component names and app ids are the whole API your code sees. statestore is Redis on your laptop and Cosmos DB in Azure, and the handler in hop 4 does not change, because the swap happens in a YAML file on one side and a Terraform resource on the other.
Container Apps owns everything at run time that is not your code: ingress, revisions, scaling on HTTP rate or KEDA queue depth, and the sidecar itself. That last one is managed in both directions. You do not install daprd, and you do not pick its version either; the platform runs 1.16 while dapr init on your machine gives you 1.18.
Terraform owns what the AppHost saved you from writing down: the VNet, the environment, the identities and every grant, the components with their scopes. Part 7 is that tree. It deploys orders-api and nothing else: there is no inventory-api, no Front Door and no API Management in it, and this article does not add them.
Nobody owns the seams between them, and the first one shows up as soon as you put Part 7's container definition next to Part 4's ServiceDefaults. The Terraform side, trimmed to the environment variable and the probes:
container {
name = each.key
image = each.value.image
cpu = each.value.cpu
memory = each.value.memory
env {
name = "ASPNETCORE_ENVIRONMENT"
value = var.environment == "prod" ? "Production" : "Development"
}
liveness_probe {
transport = "HTTP"
port = each.value.target_port
path = "/healthz/live"
}
readiness_probe {
transport = "HTTP"
port = each.value.target_port
path = "/healthz/ready"
}
}
And the method every service calls on startup:
public static WebApplication MapDefaultEndpoints(this WebApplication app)
{
// Adding health checks endpoints to applications in non-development environments has security implications.
// See https://aka.ms/aspire/healthchecks for details before enabling these endpoints in non-development environments.
if (app.Environment.IsDevelopment())
{
// All health checks must pass for app to be considered ready to accept traffic after starting
app.MapHealthChecks(HealthEndpointPath);
// Only health checks tagged with the "live" tag must pass for app to be considered alive
app.MapHealthChecks(AlivenessEndpointPath, new HealthCheckOptions
{
Predicate = r => r.Tags.Contains("live")
});
}
return app;
}
Neither file is wrong on its own terms, and together they never pass a probe. In prod the app runs as Production and maps no health endpoint at all. In every other environment it runs as Development and maps /health and /alive, which are not the paths Terraform probes either. Container Apps counts anything from 200 to 399 as healthy, so a 404 fails both probes in every environment: liveness restarts the container, readiness keeps the revision out of rotation, and the first sign is a revision that never goes healthy while nothing in the app's own logs mentions a health check. Part 4 warned that the endpoints disappear outside Development; Part 7 probed /healthz/*, the shape Dapr's app health default and most Kubernetes examples use, and no part checked one against the other. The fix is in the operations section, because it is also the health strategy.
The second seam is telemetry, and it is worse because it fails quietly.
Operating it: traces, health checks and alerts
Traces: one pipeline, and the setting that splits it
The app side needs no change. ServiceDefaults already decides where telemetry goes from one environment variable:
private static TBuilder AddOpenTelemetryExporters<TBuilder>(this TBuilder builder) where TBuilder : IHostApplicationBuilder
{
var useOtlpExporter = !string.IsNullOrWhiteSpace(builder.Configuration["OTEL_EXPORTER_OTLP_ENDPOINT"]);
if (useOtlpExporter)
{
builder.Services.AddOpenTelemetry().UseOtlpExporter();
}
// Uncomment the following lines to enable the Azure Monitor exporter (requires the Azure.Monitor.OpenTelemetry.AspNetCore package)
//if (!string.IsNullOrEmpty(builder.Configuration["APPLICATIONINSIGHTS_CONNECTION_STRING"]))
//{
// builder.Services.AddOpenTelemetry()
// .UseAzureMonitor();
//}
return builder;
}
Locally the AppHost sets OTEL_EXPORTER_OTLP_ENDPOINT to the dashboard. In Azure, when the environment has the managed OpenTelemetry agent switched on, Container Apps injects the same variable pointing at the agent, plus OTEL_EXPORTER_OTLP_PROTOCOL=grpc and OTEL_RESOURCE_ATTRIBUTES. Turning on tracing in production is an environment change, not a code change, and the commented-out UseAzureMonitor() should stay commented out: with both active, every span and log line is exported twice, once directly and once through the agent.
The environment side is the part no sample carries. Part 7's environment resource sets logs_destination = "log-analytics" and nothing about traces, so as deployed, the spans never leave the replica. This fragment is not in the companion repo; the property names are checked against the 2026-07-01 template reference:
resource env 'Microsoft.App/managedEnvironments@2026-07-01' = {
name: 'cae-orders-prod'
location: location
properties: {
appInsightsConfiguration: {
connectionString: appi.properties.ConnectionString
}
openTelemetryConfiguration: {
tracesConfiguration: {
destinations: [ 'appInsights' ]
includeDapr: true
}
logsConfiguration: {
destinations: [ 'appInsights' ]
}
}
// no daprAIConnectionString, no daprAIInstrumentationKey
}
}
includeDapr: true puts the daprd spans on the same path as the app's spans, so the hop 4 call shows up as one trace with the sidecar in the middle instead of two traces with a gap. The comment at the bottom is the line that matters.
Part 5 set daprAIConnectionString next to includeDapr: true, and that combination splits the pipeline. When either daprAIConnectionString or daprAIInstrumentationKey is set, Container Apps sends the Dapr spans straight to Application Insights and skips the agent, whatever includeDapr says. Your app spans go through the agent to its destinations, the sidecar spans go somewhere else, and if the agent's destination is anything but that same Application Insights resource, the sidecar hop is missing from the trace. This is not in the Learn docs. It is confirmed in a Microsoft Q&A thread by an Azure Container Apps engineer answering for the product group, who gives the fix as setting both properties to null. The platform also enforces a version of the rule on its own: an environment that already has appInsightsConfiguration rejects daprAIConnectionString with "DaprAIConnectionString can not be set when AppInsightsConfiguration has been set" (azure-container-apps#1771), and in that same issue the product team confirms that the CLI has no supported way to clear the value once it is set.
That is also where Part 7's Terraform walks into it. azurerm_container_app_environment in azurerm 5.6.0 has dapr_application_insights_connection_string, marked ForceNew in the provider source, and nothing at all for appInsightsConfiguration or openTelemetryConfiguration; native support is still an open request (terraform-provider-azurerm#28217). So the one telemetry knob the provider offers is the legacy one that bypasses the agent, and taking it back out replaces the environment and every app and component inside it. Leave that variable null, and set the fragment above through azapi_update_resource against the environment's ID, which is what the Terraform tab on the Learn page does.
The agent has four limits to know before you rely on it:
- The Application Insights destination takes logs and traces, not metrics. The runtime and ASP.NET Core metrics ServiceDefaults records go nowhere unless you add an OTLP metrics destination, or skip the agent and use the Azure Monitor distro in the app instead.
- It is one replica (0.5 vCPU, 1.5 GB), not scalable, with no high availability and no health metrics of its own. It buffers for up to five minutes and then drops the oldest batches.
- Its configuration is per environment. Every app in the environment sends to the same destinations.
- You cannot set the Dapr sampling rate. Sampling lives in the Dapr Configuration spec, which Container Apps does not support. Part 5 already said to count the spans that arrive rather than trust a documented default; that advice still holds.
Health checks that exist outside Development
Given the two files above, pick one set of paths and map it everywhere. The probes already name /healthz/live and /healthz/ready, so the app moves to them. This is the fix, and it is not in the companion repo, which still has the Development-only version:
private const string HealthEndpointPath = "/healthz";
public static WebApplication MapDefaultEndpoints(this WebApplication app)
{
// Mapped in every environment: the Container Apps probes run in production. The paths are
// unreachable from outside the VNet unless an API Management operation forwards them.
var probes = app.MapGroup(HealthEndpointPath);
// Readiness: every registered check.
probes.MapHealthChecks("/ready");
// Liveness: only the "live" self check, so a dependency outage never triggers a restart.
probes.MapHealthChecks("/live", new HealthCheckOptions
{
Predicate = r => r.Tags.Contains("live")
});
return app;
}
The tracing filter in ConfigureOpenTelemetry already excludes anything under HealthEndpointPath, so with the constant changed it covers both probes, and the /alive constant and its clause can go.
Part 4's objection to mapping these outside Development was an unauthenticated endpoint that fans out to every dependency on each request. On this topology the paths are not reachable from the internet: ingress is internal and API Management forwards only the operations you define. If your API uses a wildcard operation, block /healthz in the gateway policy or apply Part 4's timeout, output cache and host filter to the group.
Then decide what goes into each check. Liveness must never touch a dependency, because restarting a container cannot fix Cosmos DB. Readiness is where a dependency check would go, and on a shared dependency I would still leave it out: a Cosmos DB brownout then fails readiness on every replica of orders-api at the same moment, and the endpoints that never touch Cosmos go down with the ones that do. The template registers only the self check, and for shared dependencies that is the right default.
Dapr can also probe your app's health, and that probe has the sharpest edge. Container Apps can have the sidecar probe the app (dapr.appHealth, off by default, path /healthz, a 3 second interval, a 500 ms timeout, a threshold of 3). When that check fails, Dapr unsubscribes from every topic and stops input bindings until it passes again. A subscriber under a heavy backlog that answers its probe in 600 ms stops consuming at exactly the moment the backlog is biggest. If you turn it on, point it at /healthz/live (the default /healthz is not mapped by the group above), keep that check free of dependencies, and raise the timeout.
Alerts worth waking up for
Container Apps publishes a short list of platform metrics, and three of them, plus two from Service Bus, cover most of what pages you. The thresholds are opinions; the metric names are not:
A metric alert from the CLI, for the scale ceiling on orders-api with the max_replicas = 30 from Part 7's production tfvars:
az monitor metrics alert create \
--name "orders-api-at-max-replicas" \
--resource-group "<rg>" \
--scopes "<orders-api resource id>" \
--condition "max Replicas >= 30" \
--window-size 15m \
--evaluation-frequency 5m \
--action "<action group id>"
The 5xx rule is the same command with --condition "total Requests > 50 where statusCodeCategory includes 5xx"; check the category values in metrics explorer before you rely on the filter.
Two of those rows hide caveats. RestartCount is cumulative per replica, so a threshold of zero keeps firing as long as a replica that restarted once is still alive; treat it as "some replica has restarted" rather than "a replica is restarting now". It also counts app containers only, excluding platform-managed sidecars, so a crash-looping daprd never moves it. And the Service Bus metrics split by EntityName, which is the topic, not the subscription: a topic with two subscribers shows one number for both.
The sidecar and the probes need log search alerts instead. These queries are not in the companion repo. They use the _CL tables because Part 7 sets logs_destination = "log-analytics"; with azure-monitor and diagnostic settings the tables and columns lose their suffixes and neither query runs unchanged. Probe failures and Dapr component errors land in the system logs:
ContainerAppSystemLogs_CL
| where TimeGenerated > ago(15m)
| where Reason_s == "Unhealthy" or Log_s has_any ("BackOff", "Error creating dapr component")
| summarize Events = count() by ContainerAppName_s, RevisionName_s, Reason_s
That query would have caught the health path mismatch on the first deploy. For the sidecar itself, filter the console logs on the container name, which is always daprd:
ContainerAppConsoleLogs_CL
| where TimeGenerated > ago(15m)
| where ContainerName_s == "daprd"
| where Log_s has_any ("level=error", "level=fatal", "\"level\":\"error\"", "\"level\":\"fatal\"")
| summarize Errors = count(), Sample = take_any(Log_s) by ContainerAppName_s, bin(TimeGenerated, 5m)
The has_any list covers both logfmt and JSON, because I have not confirmed which format the managed sidecar writes. Component initialization failures and authorization errors against Cosmos DB or Service Bus, which is what a missing grant looks like from inside the sidecar, are what this query is for.
Logs from two containers in one replica
The app and its sidecar write to the same console log table from the same replica, and the replica name is the join key: ContainerGroupName_g in the _CL tables. To read both sides of one incident in order:
ContainerAppConsoleLogs_CL
| where TimeGenerated > ago(10m)
| where ContainerAppName_s == "ca-orders-api-prod"
| project TimeGenerated, Replica = ContainerGroupName_g, ContainerName_s, Log_s
| order by Replica asc, TimeGenerated asc
Console logs are the fallback. Once the agent delivers traces, one operation is one query across the workspace-based Application Insights tables, where OperationId is the W3C trace id:
let op = "<trace id>";
union AppRequests, AppDependencies, AppTraces, AppExceptions
| where OperationId == op
| project TimeGenerated, Type, AppRoleName, Name, Message, ResultCode, DurationMs
| order by TimeGenerated asc
AppRoleName separates orders-api from inventory-api, and the dependency rows show where the time went. If the rows between the two services are missing, check the environment for a daprAIConnectionString before you check anything else.
What it costs: Container Apps vs Functions in West Europe
Every price in this section is West Europe, US dollars, pay-as-you-go list price from the Azure Retail Prices API, checked on 25 September 2026: no reservations, no savings plans, no enterprise agreement, and a 730-hour month. The region matters more than it looks. Container Apps vCPU is about 40% dearer in West Europe than in East US, while Functions is priced the same in both, so a comparison run on East US defaults flatters Container Apps.
Compute: one API, six ways to host it
The workload is one .NET API: 100 ms per request at 512 MB, about 10 ms of it on the CPU and the rest waiting on Cosmos DB and Service Bus. Requests arrive during 12 busy hours a day and not at all during the other 12. Everything scales out at 60% CPU; the Container Apps replicas are 0.5 vCPU and 1 GiB on the default HTTP rule of 10 concurrent requests; each column gets the subscription's whole monthly free grant, which in a real subscription it would share with every other app. The model is my arithmetic, not a Microsoft calculator's:
Functions Consumption wins every row, and at the bottom it is not close: $3.40 against $24.51 at 10M requests. It bills 0.05 GB-seconds per request and nothing for the gaps between requests, and its $0.20 per million executions undercuts the $0.56 per million that Container Apps charges for requests from outside the environment. The catch is the plan. Learn now titles it "Consumption plan hosting (legacy)", Linux hosting on it retires on 30 September 2028 (Linux apps still on runtime v3 stopped running on 30 September 2026), and it has no VNet integration, so it cannot sit behind the private topology this series built. The column is priced as Windows Consumption, which the retirement does not touch.
That leaves Flex as the Functions column that competes in practice, and one rule decides what Flex costs for a 100 ms API: a minimum billable execution period of 1,000 ms. Learn states the minimum and says nothing about what it applies to. The most specific statement I found is an accepted answer on Microsoft Q&A, from a community moderator rather than the product group: requests that overlap on one instance share a 1,000 ms block, and requests that do not overlap are each rounded up to 1,000 ms on their own. The table uses that reading, which makes steady but sparse traffic the worst case.
At 10M requests one instance sees a request about every 130 ms, only about half of them overlap another, and every one that starts a new block bills a full second for 100 ms of work. If Flex turns out to bill an instance's wall-clock busy time instead, the 10M cell drops from $62.03 to $18.38. The Flex column is the cheaper of the two instance sizes that fit: 512 MB up to 10M, 2,048 MB above that, where a per-instance concurrency of 16 packs more requests into each shared block.
The Container Apps columns have their own surprise: scale to zero saves almost nothing on a daytime API. The last replica waits out a 300-second cooldown before it goes, and even at 1M requests a month a request arrives every 1.3 seconds during the busy hours, so a minReplicas = 0 app runs all day and scales down only at night. Keeping one replica overnight at the idle rate costs about $8 a month, which is cheap insurance against the cold starts in the next section. Service invocation from orders-api to inventory-api adds nothing to the request line, because only requests from outside the environment are billable. The sidecar is on no meter either; the only statement about it I found is a Container Apps PM writing in 2023 that you are not billed for daprd, which is old and is not documentation.
The last column is a warning, not a recommendation, about Part 7's production tfvars. They declare a D4 workload profile with minimum_count = 1, and a dedicated profile brings the $0.10 an hour management fee with it: one D4 node plus that fee is $386.64 a month, whether or not an app is placed on it.
Flat, daytime or bursty: where the lines cross
The shape of the traffic moves the answer more than the volume does. The same model, with the traffic spread flat over 24 hours and packed into two hours a day:
Flat traffic is where fixed capacity pays off, and even there it takes a lot of traffic: a dedicated D4 node gets cheaper than Container Apps Consumption only at about 312M requests a month. Squeeze the same volume into fewer hours and the crossover moves further out, until on the bursty shape the model never reaches it, because two busy hours need so many instances at once that the fixed-capacity options scale out as well.
Functions Consumption stays cheaper than Container Apps Consumption at every volume in the flat and daytime shapes. Flex does not: under the per-request reading of the 1,000 ms rule, Container Apps with scale to zero costs less than Flex from about 2M requests a month on the daytime shape. Bursty traffic is the one shape where scale to zero earns its name, since the replica is gone 22 hours a day, and at 2M requests the whole month fits inside the Container Apps free grant. The bursty rows leave one thing out: every burst starts cold, and you pay for that in latency rather than dollars.
The floor, and the lines that grow under it
The compute column is the small number. This is what sits in front of it at 1M requests a month:
Three of those lines are there because the stack is private. Front Door Standard is $35, but it cannot use a Private Link origin. API Management Basic v2 is $150, but it has no outbound VNet integration, so it cannot reach an internal environment. Service Bus Standard is $9.81 with 13M operations included, but private endpoints need Premium. Each "keep it off the internet" decision costs roughly $300 to $670 a month, and together they are 99% of the floor. The public variant (Standard tiers throughout and a public Container Apps ingress that trusts a header instead of a network path) comes to about $215 a month; Part 6 is what that trades away.
The floor is also not flat. At 100M requests on the daytime shape, four lines grow with traffic:
- Front Door bills $1.50 per million requests plus egress: $150 for the requests and $41 for 5 KB responses.
- API Management includes 50M calls and charges $2.50 per million after that: $125.
- Cosmos DB serverless at $0.305 per million RU: a 1 KB point read costs 1 RU and a 1 KB write 5 RU with indexing off (Learn), so an order that reads once and writes once is 600M RU, or $183. At that volume serverless is the wrong model; 1,000 RU/s provisioned is $58.40.
- Log Analytics at $2.99 per GB: every kilobyte of telemetry per request is $284 a month at 100M requests. A traced order through two services and two sidecars writes several rows, not one, and the Dapr side cannot be sampled down on Container Apps (the operations section explains why).
With one kilobyte of telemetry per request, the stack at 100M requests costs about $2,490 a month plus $138 of compute. Compute is about 5% of the bill. Nothing in this section makes the migration pay for itself; if it pays, it pays where Part 1 said it would, in teams that stop deploying and rolling back together.
What did not survive contact with production
The seams earlier in this article came from reading the samples side by side. These come from the Dapr source, the Container Apps issue tracker and the places where the series' own samples would break once they ran for a month, and each one fails differently from its Functions equivalent.
The subscriber that never wakes up
inventory-api has no public ingress and nothing calls it over HTTP. Deploy it with the default scale rule, which is HTTP, and minReplicas = 0, and the first quiet night takes it to zero replicas. The next morning orders pile up in its subscription and nothing happens, because the component that would pull them is the Dapr sidecar, and the sidecar runs inside the replica that no longer exists. Learn says it without hedging: with ingress disabled and no minReplicas or custom scale rule, the app "scales to zero and has no way of starting back up". A Functions developer has never had to think about this, because the Service Bus trigger is its own scale controller.
The fix is the KEDA azure-servicebus rule from Part 5, with its own identity and a subscriptionName equal to the component's consumerID, or minReplicas = 1, which the cost model prices at about $8 a month. Even with the rule, the cold path is long: KEDA polls every 30 seconds, then the image is pulled, the app starts, the sidecar starts, and Dapr delivers nothing until the app channel is up. minReplicas = 1 does not cure everything either: one report has the first Event Hubs publish after one to two idle hours taking around 15 seconds on a warm replica, because the broker closes idle AMQP links and the sidecar reconnects with backoff. Measure first-call latency after an idle hour, not only after a deploy.
A 404 is not a retry
Part 2's subscriber, lifted from the sample:
app.MapPost("/events/order-created", OnOrderCreatedAsync)
.WithTopic(PubSubName, OrderCreatedTopic);
if (order is null)
{
logger.LogWarning("No stored order for {OrderId}; dropping the event.", orderCreated.OrderId);
return Results.NotFound();
}
To an ASP.NET Core developer, NotFound() is the natural answer for a missing order. To the sidecar, the status code is the delivery verdict, and on the Dapr version Container Apps runs today, a 404 is the one verdict that none of the safety nets see. I read the runtime and the Service Bus component source for the managed 1.16 line and for 1.17, because the docs only say that a 404 is logged and the message dropped:
On 1.16 the 404 row is lost data; 1.17 treats a 404 as DROP (dapr#9401), but Container Apps does not run 1.17. The message is completed, so Service Bus's own dead-letter queue never sees it, and the Dapr deadLetterTopic Part 2 introduced as the safety net is skipped as well. The same bug in a Functions Service Bus trigger throws, exhausts MaxDeliveryCount and parks the message where you can find it. The fix, which is not in the companion repo, is to say what you mean:
app.MapPost("/events/order-created", OnOrderCreatedAsync)
.WithTopic(new TopicOptions
{
PubsubName = PubSubName,
Name = OrderCreatedTopic,
DeadLetterTopic = "orders-deadletter"
});
// in the handler: DROP is copied to the dead-letter topic; a 404 on 1.16 is not
return Results.Ok(new { status = "DROP" });
If the order might not be written yet, return a 500 instead and let redelivery try again. Either way, a missing route produces the same 404 as a missing order, so a contract test that every route in /dapr/subscribe answers something other than 404 is worth its few lines.
Cosmos DB is not the Redis you tested against
Aspire runs statestore on a local Redis, and a local Redis answers before you can measure it. In Azure the same SaveStateAsync goes from the app to its sidecar, from the sidecar through a private endpoint to Cosmos DB, and back, and the middle leg is the one that surprises people. Dapr's Cosmos DB component is built on the Go SDK, and the Cosmos DB SDKs other than .NET and Java only connect in Gateway mode, so every state call takes the extra gateway hop that a .NET CosmosClient in Direct mode would skip. Cosmos DB's under-10 ms guarantee at the 99th percentile is stated for point operations with direct connectivity. I have no benchmark to give you for the gateway path, and a number measured on someone else's network would not be yours anyway: with includeDapr on, the sidecar's call to Cosmos DB is its own span, so the trace from the operations section shows you the figure for your stack.
Latency is half of it; the other half is what the component does with keys. The Dapr Cosmos DB component uses the state key as the partition key unless the request says otherwise, so every order is its own logical partition, and a transaction across keys fails unless they share one. That is why the handler in hop 4 passes OrderPartition.For(order.CustomerId) on every write, and why every read of that order has to pass the same metadata or look in the wrong partition and come back empty. Put that together with the previous section and you get this failure: a subscriber that reads without the partition, gets nothing, returns 404, and on 1.16 completes the message. For write-heavy containers, also exclude every path except id and partitionKey from indexing: Dapr's state API reads by key, and the default policy indexes every property of every write.
The sidecar version is not yours
dapr init on your laptop installs 1.18. Container Apps runs 1.16.4-msft.N, and on 13 August 2026 a Container Apps engineer answered the question about 1.17 with "No plan for v1.17 for now". You cannot pin the managed version and you cannot hold it back; planned maintenance windows only choose when an upgrade lands. The last forced upgrade, to 1.16, broke Service Bus delivery for messages carrying URI-like application properties, and the fix for session messages being delivered in parallel shipped in 1.17 and reaches Container Apps only if you ask for a backport through a support ticket.
The companion repo is on the wrong side of this: it references Dapr.AspNetCore 1.18.5 against a 1.16 sidecar. Three habits close most of the gap:
- Match your local runtime to the cloud with
dapr init --runtime-version, so that 1.17 and 1.18 behaviour (the 404 row above, for one) never reaches your laptop and misleads you. - Pin the Dapr .NET packages to the platform's minor version, and move them when the platform moves, not when NuGet does.
- Watch the Container Apps issue tracker (aka.ms/dapr-aca-updates) rather than the Dapr release notes, which describe a runtime you do not have.
Timers and workflows have no Dapr home
Two Functions features have no managed Dapr equivalent. The cron input binding runs inside the sidecar, so it fires once per replica and not at all at zero replicas, where a Functions timer trigger is a singleton. And Dapr Workflow, the orchestration building block that 1.18 was mostly about, is not supported on Container Apps; the request for it has been open since May 2023. Container Apps Jobs cannot use Dapr at all. That leaves a scheduled Container Apps Job with the plain Azure SDKs, the Durable Task SDKs with Durable Task Scheduler, or keeping the timer and the orchestration exactly where they are.
What .NET developers find different
The code changes less than the debugging does. A Service Bus trigger used to be one attribute; now it is an endpoint in C#, a component in Terraform, a scale rule with its own identity, and a consumerID that has to match a string in two other places. When a message goes missing, the question is no longer "what does the invocation log say" but which of four things dropped it: the app, the sidecar, the component or KEDA. You can step through the app under Aspire, but not through daprd; its enableApiLogging setting is the nearest thing to seeing the binding call. Service invocation gets no platform resiliency (Container Apps says so explicitly, and the Dapr resiliency spec does not deploy there, as Part 5 found), so retries and timeouts go back into the HttpClient pipeline you thought the sidecar owned. And ASP.NET Core asks questions the Functions host answered for you: Data Protection keys have to be shared across replicas, and health endpoints, as this article showed, have to exist outside Development.
Part 4's AppHost is what makes it learnable. With one F5 that starts both services and both sidecars and draws one trace across all of them, a team learns the sidecar model in a week. Without it, daprd stays someone else's problem until the first incident.
What stays on Functions
Not everything should move, and the cost model above is a large part of why. A queue-triggered function that validates a message and writes a row is a dozen lines and, on Consumption, costs nothing at a million executions a month. Its Container Apps equivalent is an endpoint, a component, a scale rule and an identity, a replica that bills while it waits for the scale-in cooldown, and no new capability in return. Event glue with no business logic, webhook receivers, blob triggers (which have no Dapr binding with the same semantics) and low-traffic internal endpoints all land in the same place.
Timers and orchestrations stay too, for the reasons in the previous section: a timer trigger is a singleton and a cron binding is not, and Durable Functions has a home that Dapr Workflow does not have on Container Apps yet. If what you want is one environment rather than one programming model, Azure Functions on Container Apps runs function apps inside the same Container Apps environment, next to orders-api, on the same VNet and the same logs. This series did not take that path, and it is the one I would look at first for the functions that remain; the one thing to check is that Durable Functions only autoscales there with the MSSQL or Durable Task Scheduler backends.
When to start this journey
The stack in this article is the end state, not the starting point, and I would build it again in the order the series did. Part 1 gave the reason to leave, the prerequisite pulled the business logic out of the function bodies, Part 2 and Part 3 swapped bindings for Dapr building blocks inside two ASP.NET Core services, and Part 4 kept the inner loop to one F5. Part 5 moved production traffic one endpoint at a time, Part 6 put a private edge in front, and Part 7 wrote it all down in Terraform. Each step works on its own: you can stop after Part 5 with services on Container Apps and no private edge, and pay a fraction of the floor.
Whether to start is a narrower question than any of those steps, and the answer comes from outside the compute bill:
- Start when the ceiling is organizational: two or more teams release through one function app, or through one Premium plan whose always-ready floor they share. Or when it is structural: a working set above the 4 GB a Flex instance allows (only Premium, up to 14 GB on EP3, and Dedicated clear it), dependencies that need a container, or a scale-out curve that cannot keep up with your spikes. Start sooner if a private edge is already a requirement, because then the floor above is not the migration's cost; you would pay most of it on Functions too.
- Stay when the workload is event-driven and bursty and one team owns it. Functions Consumption is the cheaper host for that shape at every volume in the tables above, and the migration costs a team two quarters it could spend on the product.
- Stay when the reason is price. The model never finds a volume where moving saves more than cents of compute against Functions Consumption, and the platform in front of it dwarfs the difference either way.
For a second opinion from the other direction, the Azure Architecture Center's Deploy microservices with Azure Container Apps and Dapr is Microsoft's own version of this stack, without the private edge.
Conclusion
In Part 1 the argument for leaving Functions was a set of limits, and the finished stack shows the limits were the easy part. The expensive decisions came where two correct pieces met, and none of them showed up in a single sample run on its own. So put each layer into production before the next one goes on.
Functions did not lose this comparison. It is cheaper on compute at every volume in the tables above, and for event glue it is still the better programming model. Container Apps wins where the questions are about who deploys what and what runs next to it, and it asks for about $1,730 a month of platform before it answers.
If you started a new Azure API tomorrow, would it go on Functions or on Container Apps?






