Conversation
Add a count to ThreadmillTaskBackend.acquire so a worker reserves up to `count` tasks in one broker round-trip, and fill a per-process buffer from a dedicated fetcher thread. The buffer defaults to 4 x threads and is tunable with --prefetch-count; 1 disables batching. - Redis acquire pops a round-robin batch atomically and advances the rotation one position per call - the fetcher is a daemon thread, stops on max_tasks, shutdown, or drain, abandons a full buffer once no consumer is left, and logs plus re-raises a fetch failure so the child exits non-zero - a worker whose consumers died is recycled instead of parking forever - document the option and its soft limits in the README
Measure dramatiq beside celery and the task backends on the same trivial echo task, pinned to one process, one worker thread and a prefetch of one message so it matches the others. The threadmill queues grew to 60,000 tasks and dramatiq keeps 5,000: the fixed cost of a cold worker start and stop is quantized to about a second, which swamped the marginal drain of a shallower queue and made the prefetch comparison unmeasurable. - per-queue depth on QueueUnderTest, recorded in the benchmark extra info - fail loudly instead of writing a negative throughput when a drain is degenerate (process mean below start mean) - chart height follows the row count, and its subtitle reports the depths actually measured - the chart plots threadmill at its default configuration only; the no-prefetch run stays a benchmark diagnostic, since on a local broker the two land within a percent of each other
dramatiq's Redis consumer polls rather than blocks: with its read-ahead window full it sleeps compute_backoff(0), a jittered 5-10 ms, so pinning it to one message in flight cost a sleep between every task and made the chart read 105/s instead of its real figure. Unpinned it reads 211/s, which is five times celery rather than twenty. Celery's pinned prefetch multiplier goes with it: its consumer blocks on Redis, so the pin measured nothing (2,145/s pinned against 2,080/s at its default over 20,000 tasks). The methodology is now one worker process and one thread, each queue at its own default read-ahead, and the chart says so.
Threadmill reserves four tasks per worker by default, so every queue that can be told now reads four ahead: celery through --prefetch-multiplier=4 and dramatiq through dramatiq_queue_prefetch=4. The prefetch buffer is no longer a comparison advantage. dramatiq reads 420 tasks/s at that rate, twice the 211 it scored at its own default of two; its polling consumer still pays a jittered 5-10 ms backoff roughly once per four messages, which the chart footnote states. The two Django backends keep reading one message at a time because their shipped workers expose no read-ahead setting: db_worker claims one task per loop and run_redis_tasks hardcodes max_messages=1. Patching a third party worker would measure the patch rather than the library, so they stay as they ship and the footnote says so.
An earlier harness in this repository ran dramatiq with a prefetch window of 128 and it led the field; a later commit dropped it because its single-threaded consumer sleeps a poll backoff between messages and could not be compared fairly against queues that read one message at a time. That window is the fix, not the problem. A shared rate of four still left dramatiq penalised: its jittered 5-10 ms backoff lands once per window, so the cost per task is inverse in the depth and a shallow window measures the poll, not the queue. Every queue that can be told now reads 128 ahead, which amortises that backoff to about 0.06 ms per task and leaves celery and threadmill blocking on Redis as they always did. dramatiq also takes its Results middleware back, which stores results the way celery does, so neither queue is measured discarding the value. Its depth rises to 20,000 like the other third-party queues, because at this rate a 5,000 task drain fits inside the one-second quantisation of the fixed start cost. dramatiq leads at 6,975 tasks/s against threadmill's 5,437, close to the 7,676 the earlier harness measured, and threadmill's own prefetch buffer is worth about nine percent over reading one at a time.
Storing results for parity with celery left keys the cleanup could not match: with the default result backend the key is a bare md5 hex, so the harness's dramatiq:* pattern missed it and the results stayed in Redis until their TTL expired. Naming the result namespace makes the key greppable and the existing pattern deletes it, so no new cleanup code is needed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Each task currently costs two Redis round-trips (
acquirethenacknowledge) on the critical path, so for fast tasks a worker thread spends most of its wall-clock waiting on the broker instead of running work. This patch reserves tasks ahead of time in one batched call, which amortizes that latency and keeps the pool busy, in the spirit of the utilization principle inCONTRIBUTING.md.Approach
One
acquiremethod, extended with a count, feeding one buffer per worker process:ThreadmillTaskBackend.acquire(*queues, count=1, timeout=None, worker="") -> list[TaskResult]keeps its name and grows acount, so there is no second entry point to learn. It waits for the first task only, fills the rest without waiting, and never returns an empty list. Like the rest of the interface class, it declares the contract and raisesNotImplementedError.acquire.luapops up tocounttasks round-robin across queues and returns an array, so a batch remains a single atomicEVAL.acquirewithcount=1behaves exactly as before.TaskPrefetcherdaemon thread per worker process fills a boundedqueue.Queue; worker threads drain it. A full buffer backpressures the fetcher, so memory is bounded by--prefetch-count, which defaults to4 x threadsand accepts1to disable batching.flowchart LR Q[(Ready queue)] -->|acquire, one round-trip for N tasks| F[TaskPrefetcher] F -->|bounded put| B[/task_buffer/] B --> W[WorkerThreads] W -->|acknowledge| R[(Results)]Failure and lifecycle behaviour
A fetch error is logged through the structured pipeline, recorded on the fetcher, and re-raised so the child exits non-zero, rather than disappearing into a daemon thread as if the queue were drained. A worker whose consumer threads all died now recycles instead of parking in
joinforever, which the design required a dedicated stop signal for.Benchmark chart
dramatiq joins the comparison. Every queue runs one process and one thread, and every queue that can be told reads 128 messages ahead, the rate an earlier harness in this repository used when dramatiq led the field:
dramatiq_queue_prefetch)prefetch_count)--prefetch-multiplier)Methodology, because the ranking is very sensitive to it and I had it wrong twice:
compute_backoff(0), a jittered 5-10 ms, so the cost per task is inverse in the read-ahead. Measured curve: 1 -> 111 tasks/s, 2 -> 220, 4 -> 434, 8 -> 838, 32 -> 2,866, 128 -> 6,975. An earlier revision of this branch pinned it to one message in flight and published 105 tasks/s; at its own default of two it read 211.Resultsmiddleware (store_results=True) so it stores results the way celery does. The earlier harness did the same, and without it dramatiq is measured while discarding the value.db_workerclaims one task per loop andrun_redis_taskshardcodesmax_messages=1. Patching a third-party worker would measure the patch rather than the library, so they stay as they ship and the chart footnote says so.Why dramatiq leads
Measured per task rather than inferred, because my first explanation was wrong. The two queues issue about the same number of Redis commands - 11.03 against 11.15, the latter including dramatiq's five-command result store - but threadmill's commands are far more expensive:
EVALThreadmill makes two full JSON passes per task: the fetch script decodes the payload, stamps
RUNNINGwith the lease and worker id, and re-encodes it, and the ack writes the whole serializedTaskResultplus a history entry, an eviction scan and a telemetry publish. dramatiq never rewrites a payload - its fetch isLPOPplusSADDper id - so its per-task server work is trivial. Redis is single-threaded, so that CPU is time every other client waits behind, which is the honest price of the lease visibility, durable results and telemetry threadmill offers and dramatiq does not. Filed as a separate optimization issue, since it is a redesign of the scripts rather than part of this branch.Trade-offs worth reviewing
lease_ttl. Reserved tasks are markedRUNNINGat fetch time, so a long queue of work ahead of a buffered task can outlive its lease and be reapedFAILEDbefore it runs. This is documented in the README; sizing guidance is included. Renewing leases is the upgrade path if dwell ever matters.threadsto the buffer size, so ordering is no longer strictly global.--max-tasksis a soft limit and can overshoot by up to the buffer, because a prefetched task is always finished.worker_idsrecords the process-level fetcher identity, not the executing thread, since the fetcher holds the lease.Deferred, filed as follow-ups
--prefetch-count(operator-controlled today; an unbounded value lets one worker take the whole backlog and blocks Redis for work proportional to the count).Testing
uv run pytest -m "not benchmark": 169 passed, 5 skipped (textual absent locally);tests/test_inspector.pywith the extra: 58 passed.--cov-branchreports no partial branch on any added line).uvx prek run --all-filesgreen, and the commit hooks pass.