Background Services and Job Processing

📖 16 min read

Some work doesn’t belong in a request. A nightly cleanup has no request to belong to, and a slow report or an email send would keep the client waiting for work it doesn’t need to wait on. ASP.NET Core runs this kind of work in hosted services, which live in the same process as the web app and share its configuration, logging, and dependency injection. The cost of that convenience is that the work lives and dies with the process, which decides when a hosted service is enough and when the work needs a separate worker or a persistent job library.

Hosted Services

A hosted service implements IHostedService, whose two methods the host calls at the edges of the app’s life. StartAsync runs when the app starts and StopAsync when it shuts down. Most background work derives from BackgroundService instead, an abstract implementation that calls one method, ExecuteAsync, when the service starts, and cancels the token passed to it when the service stops. The task ExecuteAsync returns represents the service’s whole lifetime.

public sealed class ExpiredSessionCleanup(
    IServiceScopeFactory scopeFactory,
    ILogger<ExpiredSessionCleanup> logger) : BackgroundService
{
    protected override async Task ExecuteAsync(CancellationToken stoppingToken)
    {
        using var timer = new PeriodicTimer(TimeSpan.FromMinutes(15));

        do
        {
            try
            {
                await using var scope = scopeFactory.CreateAsyncScope();
                var db = scope.ServiceProvider.GetRequiredService<AppDbContext>();

                var removed = await db.Sessions
                    .Where(s => s.ExpiresAt < DateTimeOffset.UtcNow)
                    .ExecuteDeleteAsync(stoppingToken);

                logger.LogInformation("Removed {Count} expired sessions", removed);
            }
            catch (Exception ex) when (!stoppingToken.IsCancellationRequested)
            {
                logger.LogError(ex, "Session cleanup failed; retrying next tick");
            }
        }
        while (await timer.WaitForNextTickAsync(stoppingToken));
    }
}

builder.Services.AddHostedService<ExpiredSessionCleanup>();

AddHostedService registers the service as a singleton, created once and kept for the app’s lifetime. Each piece of the sample handles one of the ways such a service goes wrong: the scope, the try/catch, and the timer, which the following sections cover in turn. PeriodicTimer is the right default for periodic async work. Its loop can’t overlap itself, and cancelling the token ends the wait, so the service stops promptly. Choosing between timers, and scheduling by wall-clock time rather than interval, are general .NET topics that apply here unchanged.

Startup Order

The host starts hosted services one at a time, in registration order, and calls every StartAsync before it configures the request pipeline and starts the server. A slow StartAsync therefore delays the app’s first request, and a failing one stops the app from starting at all. That makes StartAsync the place for short initialization the app shouldn’t run without, and nowhere else. HostOptions.ServicesStartConcurrently starts them in parallel instead.

BackgroundService.StartAsync starts ExecuteAsync and returns without waiting for it. Before .NET 10, ExecuteAsync ran synchronously until its first await, so blocking code at the top of it held up every service registered after it and the server. Since .NET 10 it runs on the thread pool, and the host moves on at once.

C4 · Dynamic

Hosted Services Across the Host's Lifetime

When two hosted services start, run, and stop, relative to the server.

The lifetime of two hosted services, A and B, registered in that order in an ASP.NET Core app At startup the host calls StartAsync on service A, then on service B, in registration order, and only then starts the server, which begins accepting requests. Each BackgroundService's ExecuteAsync starts when its StartAsync runs and keeps running while the app serves requests. When a stop signal arrives, the host raises ApplicationStopping and stops hosted services in reverse registration order. The server was registered last, so it stops first: it stops accepting new connections and drains in-flight requests. Only then is B stopped, then A, each seeing its stopping token cancelled. The host waits up to the shutdown timeout, 30 seconds by default, for all of this; work still running when the timeout expires is abandoned. Host Server Service A Service B A.StartAsync B.StartAsync then the server starts and ApplicationStarted fires ApplicationStopping ApplicationStopped accepting requests 1. drains, stops A.ExecuteAsync running 3. stops B.ExecuteAsync running 2. stops stop signal ShutdownTimeout, 30 s by default. Work still running at the end is abandoned. Services start in registration order and stop in reverse. The server is registered last, so it starts last and stops first.

An Unhandled Exception Stops the Whole App

Since .NET 6, an exception that escapes ExecuteAsync is logged, and then the host stops, taking the web app down with it. That is the default HostOptions.BackgroundServiceExceptionBehavior, StopHost. Before .NET 6 the exception vanished and the service silently stopped working. Setting the behavior to Ignore restores that, which leaves a dead service in a running app.

A service whose failures are transient, such as a database that is briefly unavailable, catches exceptions per unit of work, as the sample does, logs them, and carries on with the next tick or message. The filter checks the stopping token rather than the exception type. At shutdown the cancellation passes through, and the host treats it as a normal stop rather than a failure. An OperationCanceledException from anything else, such as an HttpClient timeout, is caught like any other error, because the host only excuses a cancellation while the app is stopping, and would otherwise stop the app over it. An exception that means the service can’t work at all, such as missing configuration, is better left to stop the host, where the platform’s restart and alerting notice it. A service that keeps running while failing every iteration is invisible without monitoring, so a health check reporting its last successful run belongs with it.

Scoped Services in a Hosted Service

A hosted service is a singleton, but much of what it needs is scoped, above all an EF Core DbContext. Injecting a scoped service into a singleton’s constructor makes it a captive dependency, one instance held for the app’s whole life. A DbContext held that way accumulates tracked entities and is used from whatever thread runs the loop. In Development, the container’s scope validation throws when it detects this. In other environments that validation is off by default, so the captive instance is quietly created.

The standard pattern, which Microsoft’s DI documentation recommends, is to inject IServiceScopeFactory and create a scope for each unit of work, such as one tick or one message. Everything resolved from the scope is disposed with it, as a request’s services are disposed at the end of the request. CreateAsyncScope with await using disposes services that implement only IAsyncDisposable, which a synchronous using would throw on.

For EF Core alone, IDbContextFactory<T>, registered with AddDbContextFactory, is an alternative. The factory is a singleton that the service can inject directly, and each CreateDbContextAsync() call returns a new context for the caller to dispose. It fits a service whose only scoped dependency is the context. A service that uses repositories or other services built on the scoped context still needs a scope, so that they all share one context per unit of work.

Periodic Work Across Instances

Every instance of the app runs every hosted service. Scaled out to three instances, the cleanup above runs three times every 15 minutes. For idempotent work that is merely wasteful. For work that sends email or charges cards, it is a bug.

There are three ways out. The work can be made idempotent, safe to run twice, or safe to run concurrently, for example by having each instance claim a batch of rows with a single UPDATE ... SET ClaimedBy = @instance WHERE ClaimedBy IS NULL before processing only the rows it claimed. It can move out of the web app into a single worker instance or a platform scheduler, such as a Kubernetes CronJob or a cloud function’s timer trigger, that starts one run per occurrence. Or it can move to a job library that coordinates instances through shared storage, which the last sections cover.

Queuing Work from Endpoints

An endpoint that starts slow work, such as generating a report or calling a slow third party, can queue it and return 202 Accepted immediately. A Channel<T> makes the queue. Endpoints write to it, and a hosted service reads from it:

builder.Services.AddSingleton(_ => Channel.CreateBounded<ReportRequest>(
    new BoundedChannelOptions(capacity: 100) { FullMode = BoundedChannelFullMode.Wait }));
builder.Services.AddHostedService<ReportWorker>();

app.MapPost("/reports/{id:int}", async (int id, ClaimsPrincipal user,
    Channel<ReportRequest> queue, CancellationToken ct) =>
{
    await queue.Writer.WriteAsync(new ReportRequest(id, user.Identity!.Name!), ct);
    return Results.Accepted($"/reports/{id}");
});

// Types go after the top-level statements, or in their own files
public sealed record ReportRequest(int ReportId, string RequestedBy);

public sealed class ReportWorker(
    Channel<ReportRequest> queue,
    IServiceScopeFactory scopeFactory,
    ILogger<ReportWorker> logger) : BackgroundService
{
    protected override async Task ExecuteAsync(CancellationToken stoppingToken)
    {
        await foreach (var request in queue.Reader.ReadAllAsync(stoppingToken))
        {
            try
            {
                await using var scope = scopeFactory.CreateAsyncScope();
                var generator = scope.ServiceProvider.GetRequiredService<ReportGenerator>();
                await generator.GenerateAsync(request, stoppingToken);
            }
            catch (Exception ex) when (!stoppingToken.IsCancellationRequested)
            {
                logger.LogError(ex, "Report {ReportId} failed", request.ReportId);
            }
        }
    }
}

The queue carries data, not work. The endpoint copies what the job needs, here the report ID and user name, into the message, because the request’s HttpContext and its scoped services, including its DbContext, are disposed as soon as the response is sent. A job that captures them in a lambda fails later with ObjectDisposedException, or worse, works in testing and fails under load. The worker creates its own scope per message for the same reason. Starting the work with Task.Run from the endpoint has the same problem, and adds that nothing waits for it at shutdown.

A bounded channel with FullMode.Wait applies backpressure. When 100 reports are waiting, the next endpoint call waits for space rather than growing memory without limit. Channel options beyond that, such as dropping items when full or allowing several readers, are general System.Threading.Channels mechanics.

The channel is memory, so everything in it is lost when the process stops, whether by deployment, a crash, or scaling in. Only one instance sees each message, too, because each instance has its own channel. That is acceptable for work a client can safely request again. Work that must happen needs a durable queue, such as a message broker or a persistent job library.

Graceful Shutdown

When the app is told to stop, by a deployment, Ctrl+C, or an orchestrator’s SIGTERM, the host fires ApplicationStopping, one of the IHostApplicationLifetime tokens, and then stops the hosted services in reverse registration order. The web server is itself a hosted service, registered after all of the app’s own, so it stops first. It stops accepting connections and drains in-flight requests before any BackgroundService hears about the shutdown. For a BackgroundService, stopping means cancelling the stoppingToken and waiting for ExecuteAsync to return. The host waits up to HostOptions.ShutdownTimeout, 30 seconds by default, for the whole sequence, drain included, and abandons whatever is still running when it expires. Slow requests at shutdown therefore leave less of the timeout for background work. ServicesStopConcurrently stops them all at once instead.

A service that passes the token to every await, as the samples do, stops within moments. One that ignores the token keeps running until the timeout, and its work is cut off wherever it happens to be. Work that can’t finish in time needs to be interruptible: process items in small transactions and mark each one done, so a restarted service picks up where the last one stopped instead of redoing or losing work. Raising the timeout buys time, but only up to the platform’s own limit. An orchestrator that kills the process after its grace period, 30 seconds by default in Kubernetes, ends it regardless of what the host wanted.

builder.Services.Configure<HostOptions>(options =>
    options.ShutdownTimeout = TimeSpan.FromSeconds(60));   // also raise the platform's grace period

Some hosts stop the process on their own schedule. IIS shuts down an app pool after 20 idle minutes by default and recycles it periodically, and Azure App Service unloads idle apps unless Always On is enabled. A hosted service there stops when no requests arrive, which suits request-driven work and defeats a nightly job.

Worker Services

When background work needs more CPU or memory than the web app can spare, has a different release cadence, or must scale separately, it moves to a worker service: a separate process built on the same generic host, without the web server. The dotnet new worker template creates the project, and the Microsoft.Extensions.Hosting.WindowsServices or .Systemd package lets it run under a service manager:

var builder = Host.CreateApplicationBuilder(args);

builder.Services.AddHostedService<ReportQueueConsumer>();   // reads report requests from a message broker
builder.Services.AddWindowsService();   // or AddSystemd() on Linux; each does nothing outside a service manager

builder.Build().Run();

The hosted services, DI, configuration, and logging are identical, so code moves between the two hosts unchanged. What changes is how work arrives. A worker has no endpoints, so it reads from a queue, a database table, or a schedule. The web app’s in-memory channel can’t reach another process, which is why the consumer above reads from a message broker instead. In exchange, a heavy job can’t slow down requests, a crash takes down only the worker, and each side scales by its own measure, such as request rate for the web app and queue length for the worker.

Persistent Jobs

An in-process queue loses work on restart, retries nothing, and shows no history. Job libraries backed by a database keep jobs across restarts and share them among instances. The two most established for .NET start from different ends. Hangfire is built around a persistent job queue with retries and scheduling added. Quartz.NET is built around a scheduler, which keeps its schedule in memory unless it is configured with a database store and clustering.

Hangfire

Hangfire, through the Hangfire.AspNetCore package and a storage package such as Hangfire.SqlServer, stores jobs in SQL Server, Redis, PostgreSQL, or other storage, and its server component, running inside the web app or a worker, fetches and executes them. A job is a method call recorded as an expression:

builder.Services.AddHangfire(config => config.UseSqlServerStorage(connectionString));
builder.Services.AddHangfireServer();

app.MapPost("/orders/{id:int}/confirm", (int id, IBackgroundJobClient jobs) =>
{
    jobs.Enqueue<OrderEmails>(emails => emails.SendConfirmationAsync(id, CancellationToken.None));
    return Results.Accepted();
});

app.Services.GetRequiredService<IRecurringJobManager>()
    .AddOrUpdate<SessionCleanup>("session-cleanup", job => job.RunAsync(CancellationToken.None), Cron.Hourly());

app.MapHangfireDashboard();

Hangfire serializes the method’s arguments into storage, so a job takes an order ID rather than an order object, and the job loads fresh data when it runs. The OrderEmails instance is created from the DI container when the job executes, with its own scope. Hangfire replaces a CancellationToken argument with one that fires when the server shuts down or the job is deleted. Beyond immediate jobs, it supports delayed jobs, recurring jobs on a cron schedule, and continuations that run after another job succeeds.

A job that throws is retried automatically, 10 times by default with growing delays, and then moves to the Failed state, where the dashboard can retry it by hand. Hangfire guarantees that a job runs at least once, not exactly once. Its docs warn that, since no distributed system detects failures perfectly, the same job can in corner cases be processed on two workers, so jobs have to be idempotent, for example by recording that the email was sent and checking before sending. A recurring job whose previous run is still going can also overlap with the next one. The [DisableConcurrentExecution] attribute prevents that on a best-effort basis.

Enqueuing has its own gap. If the endpoint commits the order and the process dies before Enqueue, the email is never sent. Enqueuing first and committing second risks the opposite. Writing the job in the same transaction as the business change closes it. The usual way is an outbox: the endpoint inserts a row describing the job into a table in the same transaction as the order, and a background service enqueues each committed row and marks it sent. Relying on Hangfire’s own enqueue to join the app’s transaction is less dependable, since whether it enlists depends on the storage provider and how the connection is shared. The dashboard shows queued, running, failed, and succeeded jobs with their exceptions, and can trigger or delete them. It allows only local requests by default, and exposing it means adding an authorization filter.

Quartz.NET

Quartz.NET separates jobs, the work, from triggers, the schedule, so one job can have several triggers and a trigger can use calendars that exclude holidays or maintenance windows:

builder.Services.AddQuartz(q =>
{
    var job = new JobKey("nightly-cleanup");
    q.AddJob<NightlyCleanupJob>(o => o.WithIdentity(job));
    q.AddTrigger(t => t.ForJob(job).WithCronSchedule("0 0 2 * * ?"));   // 02:00 every day
});
builder.Services.AddQuartzHostedService(o => o.WaitForJobsToComplete = true);

[DisallowConcurrentExecution]
public sealed class NightlyCleanupJob(AppDbContext db) : IJob
{
    public async Task Execute(IJobExecutionContext context)
    {
        await db.Sessions.Where(s => s.ExpiresAt < DateTimeOffset.UtcNow)
            .ExecuteDeleteAsync(context.CancellationToken);
    }
}

Quartz cron expressions start with a seconds field and use ? for “no specific value” in the day-of-month or day-of-week field, so they aren’t interchangeable with five-field Unix cron. [DisallowConcurrentExecution] keeps a slow run from overlapping the next trigger, and WaitForJobsToComplete makes shutdown wait for running jobs within the host’s timeout. Each run gets its own DI scope, so a job can take scoped services such as a DbContext in its constructor. A trigger whose time passes while the scheduler is down or busy has misfired, and its misfire instruction decides whether it fires once as soon as possible or skips to its next time.

Quartz doesn’t retry a failed job the way Hangfire does. A job that throws a JobExecutionException can ask to be re-fired immediately, and anything more, such as a delay between attempts, is the job’s own code. Its official dashboard package, Quartz.Dashboard, is new and still marked as a work in progress, so run history and manual control are less mature than Hangfire’s.

Since Quartz.NET 4.0, AddQuartzHostedService ships in the main Quartz package. On 3.x it comes from Quartz.Extensions.Hosting, which in 4.x is an empty package kept for compatibility.

By default Quartz keeps schedules in memory. With the ADO.NET job store and clustering enabled, several instances share one database, and each trigger fires on only one of them. If an instance dies mid-job, another re-runs the job only if the job was marked to request recovery. Clustered nodes need clocks synchronized to within a second, because they coordinate through timestamps in the database.

Choosing an Approach

Requirement Approach
Periodic work, safe to run on every instance or run by one BackgroundService with PeriodicTimer
Offload work from a request; losing it on restart is acceptable Channel<T> read by a BackgroundService
Heavy or independently scaled work A worker service fed by a durable queue
Work that must survive restarts, with retries and a dashboard Hangfire
Rich schedules: calendars, many triggers per job, misfire rules Quartz.NET with a persistent, clustered store
One run per occurrence across instances, with no library A platform scheduler such as a Kubernetes CronJob

The in-process options add no infrastructure and cost nothing to run, and the persistent ones add a database schema, polling, and something new to monitor. A job library pays for itself when lost or duplicated work has a cost, not merely because the work runs on a schedule.

Key Takeaways

  • A BackgroundService runs for the app’s lifetime in the web process. Since .NET 6, an unhandled exception from ExecuteAsync stops the whole app, so catch per unit of work and let only fatal errors escape.
  • Hosted services start in registration order before the server starts, and stop in reverse order within ShutdownTimeout, 30 seconds by default. Pass the stopping token everywhere.
  • Hosted services are singletons. Create a scope per unit of work with IServiceScopeFactory.CreateAsyncScope, or use IDbContextFactory<T> when a DbContext is the only scoped dependency.
  • Every instance runs every hosted service, so periodic work that must run once needs coordination, a single worker, a platform scheduler, or a job library.
  • Queued work carries data, never the request’s HttpContext or scoped services. An in-memory channel loses its contents on restart.
  • Hangfire keeps jobs in storage and shares them across instances. Quartz.NET does the same only with a database job store and clustering. Hangfire retries failed jobs and guarantees at-least-once execution, so its jobs must be idempotent, and its dashboard needs authorization before it is exposed.

Found this guide helpful? Share it with your team:

Share on LinkedIn