How an Unreferenced asyncio Task Let the Garbage Collector Silently Drop 1,140 Webhook Deliveries
← Back
October 3, 2026Python9 min read

How an Unreferenced asyncio Task Let the Garbage Collector Silently Drop 1,140 Webhook Deliveries

Published October 3, 20269 min read

09:14 UTC. A Zendesk ticket from the largest fulfillment partner gets escalated into #payment-ops-page: "order-fulfilled webhooks missing for three orders this morning, second time this week." The on-call engineer pulls up the webhook delivery dashboard. Success rate: 100%. Failed deliveries in the last 24 hours: 0. Every graph is green. The partner is reporting something broken, and every system on this side is reporting nothing wrong.


the setup

Order events flow through an internal queue into a Python consumer service, notifier, whose job is to call each partner's webhook URL whenever an order changes state. The consumer loop pulls a batch, and for each event, dispatches a delivery without waiting for it, because waiting would mean the next batch can't start until every slow partner endpoint in the current one responds:

notifier/consumer.py, before
async def run_consumer():
    async for batch in queue.poll():
        for event in batch:
            asyncio.create_task(deliver_webhook(event))
        await queue.ack(batch)

async def deliver_webhook(event: OrderEvent):
    try:
        async with httpx.AsyncClient(timeout=10) as client:
            resp = await client.post(event.partner_url, json=event.payload)
            resp.raise_for_status()
        log.info("webhook delivered", event_id=event.id)
    except Exception:
        log.exception("webhook delivery failed", event_id=event.id)
        await dead_letter_queue.put(event)

This looks correct. Every failure path logs. Every failure path also lands the event in a dead letter queue that a separate dashboard tracks. The team had reasoned about what happens when deliver_webhook raises. Nobody had reasoned about what happens if it never runs at all.


the scramble

First theory: the partner's endpoint is flaky and they're not seeing their own failures. Their engineer pulls their ingress logs for the three missing order IDs and comes back within the hour: zero incoming requests for any of them. Not a 500, not a timeout, not a connection reset. Nobody on this side ever knocked on their door.

Second theory: the queue dropped the messages before notifier ever saw them. Queue offsets say otherwise. Each of the three events was consumed exactly once and acked at a normal timestamp, no redelivery, no gap.

Third theory: an exception inside deliver_webhook that somehow bypasses the except block, maybe something raised during client teardown after the response is already in flight. Grepping Datadog logs for the three event IDs kills this one in under a minute, there's no log line from deliver_webhook at all for any of them. Not a failure log, not a success log. The function's own first line, the try block's entry, apparently never executed, or if it did, it never got far enough to log anything.


the hunt

That's the detail that redirects the whole investigation: not "delivery failed," but "delivery was scheduled and then nothing happened." asyncio.create_task returns a Task object immediately, before any of the coroutine's body runs. The call site in run_consumer never stores that return value anywhere. It's thrown away the instant the loop moves to the next event in the batch.

Python's own asyncio documentation calls this out directly, in language specific enough that it reads like it was written for this exact incident:

docs.python.org/3/library/asyncio-task.html
Important: Save a reference to the result of this function, to avoid
a task disappearing mid-execution. The event loop only keeps weak
references to tasks. A task that isn't referenced elsewhere may get
garbage collected at any time, even before it's done.

A ten-line reproduction confirms it locally. Create a thousand detached tasks in a tight loop, each sleeping briefly before logging, then force a collection cycle:

repro.py
import asyncio, gc

async def work(i):
    await asyncio.sleep(0.05)
    print("done", i)

async def main():
    for i in range(1000):
        asyncio.create_task(work(i))   # no reference kept
    gc.collect()
    await asyncio.sleep(0.2)

asyncio.run(main())

On a stock interpreter this consistently prints well under 1,000 lines, never all of them, and the exact count varies run to run depending on allocation pressure and when the collector runs relative to task scheduling. The task objects created with no external reference are only kept alive by a weak reference inside the loop's internal bookkeeping. Nothing about that weak reference stops the garbage collector from reclaiming a task before the event loop gets a chance to actually step it.

In notifier, each batch of events creates a new detached task per event, and the for loop's local variables get overwritten on every iteration. Under light load, the event loop schedules and runs each task before the next batch arrives, so the race rarely resolves badly. Under the batch sizes this service actually sees in production, hundreds of detached tasks exist simultaneously with nothing holding a strong reference to any of them, and the interpreter's cyclic collector runs on its own schedule, independent of the event loop's. Any task not yet picked up by the loop when a collection pass happens to run is a candidate for disappearing before its first await.


the find

The root cause isn't a bug in deliver_webhook. It's that the only reference to each delivery task was the return value of asyncio.create_task, discarded at the call site, and the event loop's weak reference to a scheduled-but-not-yet-run task is not a guarantee that task survives to run. This is documented, specific, and not something a stack trace will ever show you, because a garbage-collected task doesn't raise, it just stops existing.


the fix

The official pattern keeps a strong reference to every in-flight task in a module-level set, removing each entry only once it's actually done:

notifier/consumer.py, after
_background_tasks: set[asyncio.Task] = set()

async def run_consumer():
    async for batch in queue.poll():
        for event in batch:
            task = asyncio.create_task(deliver_webhook(event))
            _background_tasks.add(task)
            task.add_done_callback(_background_tasks.discard)
        await queue.ack(batch)

add_done_callback fires once the task completes, whether it succeeded, raised, or was cancelled, and removes it from the set at that point rather than leaking the set's size forever. The set itself is the strong reference that keeps every task alive from creation to completion, independent of whatever the collector decides to do with everything else.

The team also added the one signal that would have caught this in minutes instead of days: a gauge comparing tasks created against tasks completed per minute, exported from the same set:

notifier/metrics.py
def record_task_gauges():
    in_flight.set(len(_background_tasks))
    # alert if in_flight grows without bound relative to batch throughput

Delivery success rate alone can't catch this failure mode, because a task that never runs never reports success or failure. A gauge on how many tasks are actually in flight, versus how many the consumer has created, is the only metric that would have shown a gap.


the aftermath

1,140 Webhook deliveries silently dropped over 9 days
0.8% Of total deliveries in that window, the drop rate
0 Errors, retries, or dead-letter entries generated by any of them
6 partners Affected, none large enough individually to trip an alert on volume alone
  • asyncio.create_task with no stored reference is not "fire and forget," it's "fire and maybe." The event loop's weak reference is bookkeeping, not a lifetime guarantee.
  • A dropped task never raises, so exception logging and dead-letter queues, however carefully built, can't catch this. They only see work that actually started.
  • The real tell during the hunt was "no log line at all," not "a failure log." That absence is the signature of a task that never ran its first line, and it points away from retry logic and toward task lifecycle from the start.
  • Any service creating background tasks in a loop needs a collection holding strong references for the task's full lifetime, plus a gauge on in-flight count. Success-rate dashboards alone will stay green through this exact failure.

Nothing about the fix changed throughput. The consumer still doesn't wait on delivery before acking the next batch. The only difference is that something now holds onto the work long enough to guarantee it actually happens, instead of trusting the event loop to get to it before the collector decides otherwise.

Share this
← All Posts9 min read