MQTT Python client bench

Methodology

Generated from the harness's own check definitions and this campaign's manifest, so it cannot drift from what was enforced.

Three parties, no shared code

The client under test runs alone in its own process and uv environment. Across the broker from it is a neutral C peer (sink, source or echo). Mosquitto's $SYS counters are read fresh before and after each run. Each party has its own physical core, and the orchestrator reads CPU, memory and context switches from /proc from outside the client. Worker and peer follow one absolute CLOCK_MONOTONIC schedule: warm-up 2.0 s, a measured window of 8.0 s, then 2.0 s of drain so totals reconcile. Latency is the time from a stamp in the first 8 payload bytes to arrival, binned the same way in C and Python (16 buckets per power of two, at most 6.25 % wide); a cell's percentiles come from the merged histogram of its valid runs.

The stamp is the actual publish time, so latency is transit only. On fixed-rate points the client's messages are numbered, and message n is due at tstart + (n + 1) / rate. The worker keeps each publish's send time and, after the run, reports the schedule lag: send time minus due time, over the messages due in the window, plus the count due there that were never published. A client that falls behind its schedule, or whose awaited publishes run out of in-flight slots, shows it there even while it still holds the rate. The harness paces in 1 ms ticks, so up to 1 ms of lag is the pacer's, the same for every client; tails beyond that are the client's.

Statuses

valid
Every check passed. Only valid runs enter a median.
not_sustained
The client could not hold the fixed offer. A real finding about the client, kept out of cost and latency medians because a backlog's latency is queueing time.
invalid
The harness, the peer, the broker or the host failed. The run says nothing about the client; the campaign retries it once.

Checks

Counts from two parties must agree within 5 + 0.05% of the count.

worker_completed
The client worker connected, followed the schedule and reported its counters.
peer_completed
The C peer connected, followed the schedule and reported its counters.
source_completed
Duplex: the C source feeding the client connected, followed the schedule and reported.
broker_counters_read
A fresh $SYS reading was taken before and after the run.
broker_confirms_client_publishes
Publishes the broker received from the client lie between the client's completions and its sends (round trips subtract the echo's republishes, duplex the C source's publishes).
broker_confirms_peer_publishes
Publishes the broker received equal what the C source wrote.
broker_confirms_deliveries
Messages the broker sent equal what the receiving side counted. At receive capacity, only receiving more than was sent fails: a slow client still has data in its socket buffers when the run stops.
no_loss
QoS 1 and 2: every acknowledged publish reached its subscriber. A client that fell behind until the broker's queue overflowed is not_sustained; a loss without drops is invalid.
no_loss_inbound
Duplex, QoS 1 and 2: every publish the broker acknowledged to the C source reached the client.
payloads_intact
Every payload the C sink received had one of the lengths the client published.
callbacks_matched
Filter points: every message reached a per-filter callback, none the catch-all on_message.
properties_delivered
MQTT 5 properties: every message the C sink received still carried its user properties.
topic_alias_used
Topic alias: the bytes the broker received per client publish, net of the sink's acknowledgements, are below the payload plus half the topic, so the long topic was not sent each time.
offered_rate_held
At least 98% of the fixed offer was produced in the window.
responses_kept_up
Round trips: at least 98% of the requests were answered in the window.
client_kept_up
Receive and duplex: the client took at least 98% of the offer in the window.
source_rate_held
Duplex: the C source wrote at least 98% of its fixed offer in the window.
broker_headroom
Fixed-rate and idle points: the broker used less than 85% of its core.
host_quiet
At most 0.5 cores were busy outside the client, the peer and the broker. Enforced on comparable profiles only. The detail names the busiest other processes and the kernel time no process is charged for (softirq, irq, iowait, steal).

Flags

A flag qualifies a valid run without invalidating it.

broker_bound
Capacity point where the broker used at least 85% of its core: the rate is partly the broker's.
offer_bound
The client received the whole receive offer; its capacity is at least this rate.
broker_queue_overflow
The broker discarded QoS 1 or 2 messages a slower client could not drain.
host_noisy
The rest of the host was busy (only tolerated on non-comparable profiles).
non_comparable
Development profile: never published or compared.

Metrics

msgs/s
Messages confirmed delivered per second of the window. Higher is better.
CPU / msg
Client CPU (user + sys, all threads) per message. Lower is better.
user / msg
Client user-mode CPU per message. Lower is better.
sys / msg
Client kernel-mode CPU per message. Lower is better.
CPU
Share of one core the client used over the window. Lower is better.
peak RSS
Peak resident memory of the client during the window. Lower is better.
threads
Client threads at the end of the window. Lower is better.
ctx sw / 1k
Context switches per thousand messages. Lower is better.
connect
Connect call to CONNACK. Lower is better.
in flight at stop
Messages the broker sent that the client had not read when the run stopped. Lower is better.
p50
Median latency over every valid sample. Lower is better.
p90
90th percentile latency. Lower is better.
p99
99th percentile latency. Lower is better.
p99.9
99.9th percentile latency. Lower is better.
max
Largest latency observed. Lower is better.
lag p99
99th percentile of how late the client published against the fixed schedule. Up to 1 ms of it is the harness's pacing tick, identical for every client. Lower is better.
lag max
Latest publish against the fixed schedule. Lower is better.
receive p50
Duplex: median latency of the messages the client received. Lower is better.
receive p99
Duplex: 99th percentile latency of the messages the client received. Lower is better.

Points

Fixed-rate points offer every client the identical absolute load, so their costs and latencies compare across all libraries. A client that cannot hold the offer is reported as not sustained rather than timed on its backlog.

pointkindQoSpayloadloadprotocolquestion
pub_qos0_maxpub0256 BcapacityMQTTv311How many QoS 0 messages can the client publish per second?
pub_qos1_maxpub1256 Bcapacity, 64 in flightMQTTv311How many QoS 1 messages (PUBACK received) per second, 64 in flight?
pub_qos1_fixedpub1256 B2,000 msgs/sMQTTv311At 2,000 QoS 1 msgs/s, what does publishing cost and how long until delivery?
sub_qos0_maxsub0256 BcapacityMQTTv311How many QoS 0 messages can the client receive per second?
sub_qos1_fixedsub1256 B2,000 msgs/sMQTTv311At 2,000 QoS 1 msgs/s, what does receiving cost and how late do messages arrive?
rtt_qos1_fixedrtt1256 B1,000 requests/sMQTTv311At 1,000 requests/s against a neutral echo, what is the application round trip?
pub_16k_fixedpub116 KiB1,000 msgs/sMQTTv311At 1,000 QoS 1 msgs/s of 16 KiB, what do large payloads cost?
idle_connectidle0256 BidleMQTTv311How long does connecting take, and what does an idle connection cost?
pub_qos1_fixed_v5pub1256 B2,000 msgs/sMQTTv5At 2,000 QoS 1 msgs/s, what does publishing cost and how long until delivery?
sub_qos1_fixed_v5sub1256 B2,000 msgs/sMQTTv5At 2,000 QoS 1 msgs/s, what does receiving cost and how late do messages arrive?
rtt_qos1_fixed_v5rtt1256 B1,000 requests/sMQTTv5At 1,000 requests/s against a neutral echo, what is the application round trip?
sub_qos1_maxsub1256 BcapacityMQTTv311How many QoS 1 messages can the client receive per second?
pub_qos1_max_v5pub1256 Bcapacity, 64 in flightMQTTv5QoS 1 publish capacity over MQTT 5.
sub_qos1_max_v5sub1256 BcapacityMQTTv5QoS 1 receive capacity over MQTT 5.
pub_qos1_max_16kpub116 KiBcapacity, 64 in flightMQTTv311QoS 1 publish capacity with 16 KiB payloads.
pub_qos1_fixed_500pub1256 B500 msgs/sMQTTv311At 500 QoS 1 msgs/s, what does publishing cost?
pub_qos1_fixed_5kpub1256 B5,000 msgs/sMQTTv311At 5,000 QoS 1 msgs/s, what does publishing cost?
sub_qos1_fixed_500sub1256 B500 msgs/sMQTTv311At 500 QoS 1 msgs/s, what does receiving cost?
sub_qos1_fixed_5ksub1256 B5,000 msgs/sMQTTv311At 5,000 QoS 1 msgs/s, what does receiving cost?
rtt_qos0_fixedrtt0256 B1,000 requests/sMQTTv311Round trip at 1,000 req/s with QoS 0.
pub_qos2_maxpub2256 Bcapacity, 64 in flightMQTTv311How many QoS 2 messages (PUBCOMP received) per second, 64 in flight?
pub_qos2_fixedpub2256 B2,000 msgs/sMQTTv311At 2,000 QoS 2 msgs/s, what does publishing cost?
sub_qos2_fixedsub2256 B2,000 msgs/sMQTTv311At 2,000 QoS 2 msgs/s, what does receiving cost?
pub_fanout_fixedpub1256 B2,000 msgs/sMQTTv311At 2,000 QoS 1 msgs/s spread over 1,000 topics, what does publishing cost?
sub_fanin_fixedsub1256 B2,000 msgs/sMQTTv311At 2,000 QoS 1 msgs/s from 1,000 topics through one wildcard, what does receiving cost?
sub_filters_fixedsub1256 B2,000 msgs/sMQTTv311The same, dispatched to 100 per-filter callbacks (message_callback_add)?
duplex_qos1_fixedduplex1256 B1,000 msgs/s each wayMQTTv311Publishing and receiving 1,000 QoS 1 msgs/s each at once, what does it cost?
pub_qos1_fixed_tlspub1256 B2,000 msgs/sMQTTv311At 2,000 QoS 1 msgs/s over TLS, what does publishing cost?
sub_qos1_fixed_tlssub1256 B2,000 msgs/sMQTTv311At 2,000 QoS 1 msgs/s over TLS, what does receiving cost?
idle_connect_tlsidle0256 BidleMQTTv311How long does a TLS connect take, and what does an idle TLS connection cost?
pub_qos1_fixed_v5_propspub1256 B2,000 msgs/sMQTTv5At 2,000 QoS 1 msgs/s with four PUBLISH properties, what does publishing cost?
sub_qos1_fixed_v5_propssub1256 B2,000 msgs/sMQTTv5At 2,000 QoS 1 msgs/s with four PUBLISH properties, what does receiving cost?
pub_qos1_fixed_v5_aliaspub1256 B2,000 msgs/sMQTTv5At 2,000 QoS 1 msgs/s on a 200-byte topic sent as a topic alias, what does publishing cost?
sub_qos1_max_v5_rm16sub1256 BcapacityMQTTv5QoS 1 receive capacity when the client allows only 16 unacknowledged deliveries (Receive Maximum).
pub_64k_fixedpub164 KiB500 msgs/sMQTTv311At 500 QoS 1 msgs/s of 64 KiB, what does publishing cost?
pub_1m_fixedpub11 MiB50 msgs/sMQTTv311At 50 QoS 1 msgs/s of 1 MiB, what does publishing cost?
pub_rl_boundariespub1packets of 127 B, 128 B, 16383 B, 16 KiB, 2097151 B, 2 MiB60 msgs/sMQTTv311Are packets on each side of a remaining-length step (127/128 B, 16 KiB, 2 MiB) all delivered intact?

Sessions

Session 1

host
yoch-HP · Intel(R) Core(TM) i7-3770 CPU @ 3.40GHz · kernel 6.8.0-142-lowlatency · governor performance
cores
broker 0,4, sut 1,5, peer 2,6, orch 3,7
broker
eclipse-mosquitto:2.1.2-alpine@sha256:38c0da4f2ef84284d47b3b3eeea1cb3bdeabe81ee10caf0cd5c5ff61ee3ea408 · config b2b573858eaf
harness
c857af525ae96ffe8990132d3f41a0d64ae1e734 · peer 7a1328f1e2b06d35
harness floor
publish_sync 537 ns, publish_nowait 506 ns, publish_awaited 387 ns, receive_stamped 383 ns, payload_stamp 283 ns, python_loop 18 ns · worker RSS floor 20504 KiB · budget 1000 ns
broker C→C ceiling
78,749 msgs/s
receive offer
70,874 msgs/s

Raw data

Every run of this campaign, with its counts, checks and histograms, as the harness wrote it.