How an Unreferenced asyncio Task Let the Garbage Collector Silently Drop 1,140 Webhook Deliveries
09:14 UTC. A Zendesk ticket from the largest fulfillment partner gets escalated into
#payment-ops-page: "order-fulfilled webhooks missing for three orders this morning,
second time this week." The on-call engineer pulls up the webhook delivery dashboard. Success
rate: 100%. Failed deliveries in the last 24 hours: 0. Every graph is green. The partner is
reporting something broken, and every system on this side is reporting nothing wrong.
the setup
Order events flow through an internal queue into a Python consumer service,
notifier, whose job is to call each partner's webhook URL whenever an order changes
state. The consumer loop pulls a batch, and for each event, dispatches a delivery without
waiting for it, because waiting would mean the next batch can't start until every slow partner
endpoint in the current one responds:
async def run_consumer():
async for batch in queue.poll():
for event in batch:
asyncio.create_task(deliver_webhook(event))
await queue.ack(batch)
async def deliver_webhook(event: OrderEvent):
try:
async with httpx.AsyncClient(timeout=10) as client:
resp = await client.post(event.partner_url, json=event.payload)
resp.raise_for_status()
log.info("webhook delivered", event_id=event.id)
except Exception:
log.exception("webhook delivery failed", event_id=event.id)
await dead_letter_queue.put(event)
This looks correct. Every failure path logs. Every failure path also lands the event in a dead
letter queue that a separate dashboard tracks. The team had reasoned about what happens when
deliver_webhook raises. Nobody had reasoned about what happens if it never runs at
all.
the scramble
First theory: the partner's endpoint is flaky and they're not seeing their own failures. Their engineer pulls their ingress logs for the three missing order IDs and comes back within the hour: zero incoming requests for any of them. Not a 500, not a timeout, not a connection reset. Nobody on this side ever knocked on their door.
Second theory: the queue dropped the messages before notifier ever saw them. Queue
offsets say otherwise. Each of the three events was consumed exactly once and acked at a normal
timestamp, no redelivery, no gap.
Third theory: an exception inside deliver_webhook that somehow bypasses the
except block, maybe something raised during client teardown after the response is
already in flight. Grepping Datadog logs for the three event IDs kills this one in under a
minute, there's no log line from deliver_webhook at all for any of them. Not a
failure log, not a success log. The function's own first line, the try block's
entry, apparently never executed, or if it did, it never got far enough to log anything.
the hunt
That's the detail that redirects the whole investigation: not "delivery failed," but
"delivery was scheduled and then nothing happened." asyncio.create_task returns a
Task object immediately, before any of the coroutine's body runs. The call site in
run_consumer never stores that return value anywhere. It's thrown away the instant
the loop moves to the next event in the batch.
Python's own asyncio documentation calls this out directly, in language specific enough that it reads like it was written for this exact incident:
Important: Save a reference to the result of this function, to avoid
a task disappearing mid-execution. The event loop only keeps weak
references to tasks. A task that isn't referenced elsewhere may get
garbage collected at any time, even before it's done.
A ten-line reproduction confirms it locally. Create a thousand detached tasks in a tight loop, each sleeping briefly before logging, then force a collection cycle:
import asyncio, gc
async def work(i):
await asyncio.sleep(0.05)
print("done", i)
async def main():
for i in range(1000):
asyncio.create_task(work(i)) # no reference kept
gc.collect()
await asyncio.sleep(0.2)
asyncio.run(main())
On a stock interpreter this consistently prints well under 1,000 lines, never all of them, and the exact count varies run to run depending on allocation pressure and when the collector runs relative to task scheduling. The task objects created with no external reference are only kept alive by a weak reference inside the loop's internal bookkeeping. Nothing about that weak reference stops the garbage collector from reclaiming a task before the event loop gets a chance to actually step it.
In notifier, each batch of events creates a new detached task per event, and the
for loop's local variables get overwritten on every iteration. Under light load, the event loop
schedules and runs each task before the next batch arrives, so the race rarely resolves badly.
Under the batch sizes this service actually sees in production, hundreds of detached tasks exist
simultaneously with nothing holding a strong reference to any of them, and the interpreter's
cyclic collector runs on its own schedule, independent of the event loop's. Any task not yet
picked up by the loop when a collection pass happens to run is a candidate for disappearing
before its first await.
the find
The root cause isn't a bug in deliver_webhook. It's that the only reference to each
delivery task was the return value of asyncio.create_task, discarded at the call
site, and the event loop's weak reference to a scheduled-but-not-yet-run task is not a guarantee
that task survives to run. This is documented, specific, and not something a stack trace will
ever show you, because a garbage-collected task doesn't raise, it just stops existing.
the fix
The official pattern keeps a strong reference to every in-flight task in a module-level set, removing each entry only once it's actually done:
_background_tasks: set[asyncio.Task] = set()
async def run_consumer():
async for batch in queue.poll():
for event in batch:
task = asyncio.create_task(deliver_webhook(event))
_background_tasks.add(task)
task.add_done_callback(_background_tasks.discard)
await queue.ack(batch)
add_done_callback fires once the task completes, whether it succeeded, raised, or
was cancelled, and removes it from the set at that point rather than leaking the set's size
forever. The set itself is the strong reference that keeps every task alive from creation to
completion, independent of whatever the collector decides to do with everything else.
The team also added the one signal that would have caught this in minutes instead of days: a gauge comparing tasks created against tasks completed per minute, exported from the same set:
def record_task_gauges():
in_flight.set(len(_background_tasks))
# alert if in_flight grows without bound relative to batch throughput
Delivery success rate alone can't catch this failure mode, because a task that never runs never reports success or failure. A gauge on how many tasks are actually in flight, versus how many the consumer has created, is the only metric that would have shown a gap.
the aftermath
-
asyncio.create_taskwith no stored reference is not "fire and forget," it's "fire and maybe." The event loop's weak reference is bookkeeping, not a lifetime guarantee. - A dropped task never raises, so exception logging and dead-letter queues, however carefully built, can't catch this. They only see work that actually started.
- The real tell during the hunt was "no log line at all," not "a failure log." That absence is the signature of a task that never ran its first line, and it points away from retry logic and toward task lifecycle from the start.
- Any service creating background tasks in a loop needs a collection holding strong references for the task's full lifetime, plus a gauge on in-flight count. Success-rate dashboards alone will stay green through this exact failure.
Nothing about the fix changed throughput. The consumer still doesn't wait on delivery before acking the next batch. The only difference is that something now holds onto the work long enough to guarantee it actually happens, instead of trusting the event loop to get to it before the collector decides otherwise.