Prefect vs Airflow vs Temporal: three ways to decide what runs next

Swati Ahuja9 min read

prefect vs airflow vs temporal · airflow vs temporal · prefect vs airflow · workflow orchestration architecture · durable execution · airflow scheduler internals

On this page

I went into this expecting three versions of the same product, because that is what Prefect, Airflow and Temporal look like from their feature lists. Then I read the part of each codebase that answers one question, "what runs next, and who is allowed to say so", and got three answers different enough that the choice mostly follows from which answer fits your work.

Engineers like a system with one obvious place where decisions happen. All three have one, and it sits in a different place in each:

  • Airflow runs a scheduler loop. Every pass, it reads the database, locks the rows it wants to change, moves task instances forward and hands them to an executor.
  • Prefect lets your code report what it is doing ("I'm running", "I failed") and runs each report through an ordered list of server-side rules that can accept it, reject it or swap it for a different state.
  • Temporal keeps an append-only history of events per workflow. Your worker sends back commands, the server turns them into events, and the current state is whatever you get by applying those events in order.

The numbers come first, with one definition. repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. The pages show it to one decimal, so that is what I quote.

PrefectAirflowTemporal
LanguagePython (+ TypeScript UI)Python (+ TypeScript UI)Go
Code health7.66.96.7
Maintainability7.97.77.1
Files4,83910,2733,156
Lines of code (non-blank, non-comment)814,1741,538,379755,339
Symbols (functions, classes, methods)36,09690,08040,520
Hotspots73693300
Indexed17 Jul 202619 Aug 20266 Oct 2026

Scroll the table sideways to see every column.

A hotspot here is a file that changes often and is also complex, which is usually where a team spends its maintenance time.

Airflow: a loop that locks rows

The centre of Airflow is SchedulerJobRunner in airflow-core/src/airflow/jobs/scheduler_job_runner.py, a file of about 4,300 lines at the indexed commit. _run_scheduler_loop calls _do_scheduling over and over, and the docstring on _do_scheduling is one of the clearest pieces of design writing I have read in an open-source repo. It gives you the whole model in three steps:

  1. Create any DAG runs that are due (10 DAGs per loop by default, so one scheduler doesn't spend its whole pass creating runs).
  2. Pick the "next oldest" running DAG runs (20 by default) and try to move their task instances forward.
  3. In a critical section, queue task instances and send them to the executor.

What I wanted to know was how several schedulers run at once without stepping on each other. They never talk to each other at all, and they coordinate through the database instead. _critical_section_enqueue_task_instances does a SELECT ... FOR UPDATE on the pool table, so only one scheduler at a time can be in that section.

Elsewhere the file uses with_row_locks(..., skip_locked=True) again and again: if another scheduler already holds a row, skip it and do other work. On MySQL 5.x and MariaDB, which can't skip locked rows, the docstring says the other schedulers simply wait. It is a very polite form of concurrency.

For batch pipelines this is a reasonable design. The database is the source of truth, the schedulers are interchangeable workers, and you scale out by adding schedulers. The cost is that "what runs next" is always a database query, so your scheduling speed is bounded by how fast that database answers.

Full module map: Airflow architecture on repowise.

Most of Airflow is providers

Airflow is the biggest repo here by a wide margin, and the scheduler accounts for very little of that. Of its 10,273 files, about 5,900 sit in providers/, the integration packages for Google, AWS, Databricks and everyone else. On main today there are 109 provider.yaml files under providers/, one per provider package, while airflow-core (the scheduler, API and UI) is around 2,400 files and the new task-sdk is about 260.

The longest function in the whole repo shows this well. It is get_provider_info in the Google provider, at 1,703 lines, and it is one big generated dictionary: the file starts with "THIS FILE IS AUTOMATICALLY GENERATED AND WILL BE OVERWRITTEN!". The largest file is a generated OpenAPI spec. So when someone says "Airflow is huge", the accurate reply is that the ecosystem is huge and the scheduler is a few thousand lines you can read in an afternoon.

Prefect: rules that run on every state change

Prefect starts from your Python code. You put @flow and @task on functions, and as they run, the client tells the server about each state change: pending, running, failed, completed. The server records each change and also runs it through a policy before it accepts it.

The policies live in src/prefect/server/orchestration/core_policy.py. CoreTaskPolicy.priority() returns an ordered list of rule classes: CacheRetrieval, HandleTaskTerminalStateTransitions, PreventRunningTasksFromStoppedFlows, SecureTaskConcurrencySlots, CopyScheduledTime, WaitForScheduledTime, RetryFailedTasks, and more. CoreFlowPolicy has its own list of about 17, covering concurrency limits, pausing and resuming, late runs and retries. Each rule says which state it applies from and to, and has a before_transition hook.

Retries are the clearest example of how this differs from Airflow. RetryFailedTasks applies from RUNNING to FAILED, and if the task still has retries left, the rule rejects the failed state and hands back AwaitingRetry with a scheduled time, including optional jitter. The client asked to go to "failed" and the server answered "no, you're waiting to retry at 14:02". Nothing in a loop polled for failed tasks, because the decision happened at the moment the state changed.

I like this design because each behaviour is a small class you can read on its own. The trade-off is ordering. The comment on SecureTaskConcurrencySlots ("retrieve cached states even if slots are full") tells you the position in the list matters, and with 30+ rules in one 2,051-line file, the order is part of the contract.

Full module map: Prefect architecture on repowise.

Prefect's longest function has the same flavour as Airflow's. It is upgrade at 819 lines, inside the initial SQLite migration from January 2022: a schema written out table by table that nobody is meant to edit. The largest file is a test mock in the new UI.

Temporal: an event log you can replay

Temporal is the one that changes how you write your code as well as how it is scheduled. Your workflow function runs on your worker, and each time it reaches a point where it needs the outside world (call an activity, start a timer, start a child workflow), it sends a command back to the server. The server records what happened as events in that workflow's history. If your worker crashes, a new one replays the history and your function picks up where it was, as if it had never stopped, and that is what "durable execution" means.

On the server side, two files are worth knowing.

service/history/api/respondworkflowtaskcompleted/workflow_task_completed_handler.go has handleCommand, a switch over 15 command types: schedule activity, start timer, complete, fail, continue-as-new, signal another workflow, and so on. This is where "what runs next" is decided, and the workflow code decides it by what it asks for. No scheduler is looking at a graph.

service/history/workflow/mutable_state_impl.go is where those commands become state, and at 10,139 lines it is the single biggest piece of logic in any of the three repos. It follows one pattern throughout. AddActivityTaskScheduledEvent checks the workflow can still change, rejects a duplicate activity ID, asks the history builder for a new event, then calls ApplyActivityTaskScheduledEvent to update the in-memory state, and then generates the tasks that will actually dispatch the activity.

The file has 60 Add... methods and 50 Apply... methods. "Add" writes a new event and "Apply" changes state from an event. Keeping the two separate is what makes replay possible, because when the server rebuilds state from history, it only calls the "Apply" side.

Full module map: Temporal architecture on repowise. The biggest modules are common (about 1,170 files), service (about 870) and chasm (about 220), a newer internal layer for composing state machines.

Temporal's lower score

Temporal's 6.7 is the lowest of the three, a little below Airflow's 6.9 and further below Prefect's 7.6, and it would be easy to read that as "messier code". I think the score mostly reflects how much hard work the server does for itself.

Airflow and Prefect both hand the hard part of their state to a relational database, so their correctness rests on Postgres row locks and transactions. Temporal carries much more of that weight itself: event histories, sharding, replication between clusters, and task queues with their own partitioning. That is why a file like mutable_state_impl.go exists at all. A system that owns its own consistency rules ends up with long, branchy functions that touch many fields, and our health score marks long and branchy code down.

It also has 300 hotspots, files that change often and are complex. I would expect that from a server that is still adding features to the same core: Temporal's Standalone Activities, which run an activity as a job without a workflow around it, went into public preview in May 2026 and became generally available on 15 September 2026.

The same reasoning works in the other direction. Prefect's 7.6 partly reflects that a big share of the repo is a React UI and integration packages, which tend to score well. Airflow's 6.9 averages a scheduler with about 100 provider packages of mostly straightforward hook and operator classes.

Picking one

My take, from the code more than the docs:

  • If your work is batch pipelines on a schedule and you want a huge library of ready-made connectors, Airflow's model fits. The scheduler loop is simple to reason about and the database is your audit log.
  • If your work is Python-first and you want retries, caching and concurrency limits to behave the same everywhere, Prefect's rule list is the clearest of the three to read and extend.
  • If your work is long-running business processes where losing a step is not acceptable (payments, provisioning, anything that waits days for a human), Temporal's event history is the right tool, and the complexity in its server is the price of that guarantee. You pay a little of it in your own code too, because workflow functions must be deterministic so replay gives the same answer.

Limits of these numbers

  • Health scores compare the shape of the code. They say nothing direct about reliability in production, and a 6.7 server can be more reliable than a 7.6 one.
  • The three snapshots are from different dates: Prefect in July, Airflow in August, Temporal today. Code moves fast in all three.
  • Size counts everything in the repo, including UIs, tests, docs and generated files. That is why I name the largest file and function instead of leaning on totals.
  • Code links and the line counts I give for individual files are from the same commits the scores come from (Prefect b749a67, Airflow 9c8d751, Temporal a2de6d8), so they may differ from main today.
  • Prefect's index only saw its last 500 commits, which can move its score.
  • I didn't benchmark throughput, so nothing here tells you which is faster for your workload.

How we measured

Each repo was indexed by repowise: it parses every file, builds the import graph and reads the git history. repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. Numbers come from each repo's latest ready snapshot, and the code walk-through comes from reading the source on GitHub.

I did all of this by opening files and scrolling, which is a slow way to ask "where does Temporal decide to retry an activity?". You can connect any repowise repo page to Claude or ChatGPT as an MCP server and ask it that question about these three codebases, or your own. Connect it to Claude or ChatGPT.

Three codebases gave three answers to who is allowed to say what runs next, and each answer is written down in a file you can read.