QuestDB's QWP under the hood: Gorilla timestamp compression
How QWP, QuestDB's binary wire protocol, uses Gorilla delta-of-delta encoding to send a timestamp in as little as one bit, worked through on real ticks.
I've already written about QWP being faster than ILP for ingestion and about it streaming query results into Arrow. A TSBS row is about 347 bytes as line protocol text and about 97 bytes as QWP, so the same link carries roughly three and a half times as many rows.
It's easy to guess that sending data column by column is part of it. Values in a column all have the same type, they're often similar, and they often repeat, which is what any compressor wants. But if you want to see where the bytes go rather than guess, a big part of the saving comes from the encodings that run on top of that layout. This is the first of two posts about them, and it covers the one we use for timestamps: Gorilla encoding.
What a timestamp costs on the wire
Most QuestDB tables have a designated timestamp, and WAL tables, the default, require one. In its most literal form that timestamp is an 8-byte integer per row: 64 bits, whether your data arrives once a day or a thousand times a second.
The values are almost entirely predictable and they still cost the full 64 bits each. Three ticks 100 milliseconds apart cost the same as three random dates from three different centuries.
Storing the delta instead of the timestamp
The obvious first step is to send the difference instead of the value. One timestamp goes in full, the reference, and every row after it is the distance from the previous one. Those numbers are much smaller, so they fit in a narrower field.
Take a few rows of real crypto trades, timestamps in microseconds:
| Timestamp | Stored as |
|---|---|
| 09:55:57.305000 | reference, 64 bits |
| 09:55:57.312000 | +7,000 |
| 09:55:57.315000 | +3,000 |
| 09:55:57.316000 | +1,000 |
| 09:55:57.316000 | +0 |
Four small numbers instead of four full epoch timestamps. If the delta field is 32 bits wide it covers gaps of up to about 35 minutes at microsecond resolution, and any longer pause means writing a new reference in full and starting again.
This is already an encoding, and on a feed that never stops it works well. What it doesn't use is one of the most common patterns in time-series data: points arriving at regular intervals, where the gap between one point and the next is the same, or very close to it, every time.
Where Gorilla comes from: Facebook's in-memory TSDB
The encoding we chose for timestamps in QWP isn't ours. It comes from Gorilla, the in-memory time-series database Facebook built to monitor its own infrastructure, described in Gorilla: A Fast, Scalable, In-Memory Time Series Database by Pelkonen et al. at VLDB 2015.
Their problem was storing monitoring data at a scale where memory was the constraint: billions of time series and millions of data points per second, kept in RAM so queries could answer in under a millisecond. To fit that, they compressed each data point in two parts, the timestamp and the floating point value, and got their time series down to an average of 1.37 bytes per point, a 12x reduction. The abstract quotes 10x for the storage footprint of the system as a whole.
QWP takes only the timestamp half. Floating point values go on the wire at their native width, and the XOR scheme the paper uses for them isn't part of our protocol.
Delta of delta: storing how much the gap changed
A feed that samples every 100 milliseconds writes the same delta, 100,000 microseconds, on every single row. Gorilla stores how much that gap changed instead of the gap itself: the delta of the delta, or DoD.
On a steady cadence the gap doesn't change, so the DoD is zero on every row, and zero is cheap to encode. On a feed that drifts or jitters, the DoD is a small number around zero rather than a large one. In Facebook's own data, about 96% of timestamps came out as a zero.
How the decoder knows where each value ends
These numbers have different sizes, and they're packed into one continuous stream of bits with nothing between them. The decoder needs to know how many bits belong to each value, so every value starts with a short label that says how big the number after it is:
The decoder reads bit by bit. A 0 means the gap didn't change and there's
nothing else to read for that value. A 1 means the label continues: the number
of ones before the first zero says which size follows, and after four ones the
decoder stops counting and reads a 32-bit value.
The ranges are one off from the paper's, which uses -63 to 64 and so on. QWP uses the natural range of a signed integer of each width. The other difference is the start of the stream: the paper stores a two-hour block header and a 14-bit first delta, while QWP writes the first two timestamps of each column in full and starts the delta of delta at the third.
Here's a stream with a steady feed and some jitter, with the decoder walking through it:
What Gorilla costs on real trade timestamps
The diagram below is eighteen consecutive rows of the trades table on
demo.questdb.com, timestamps only, as the database
returned them. This is live crypto trade data from OKX, one TIMESTAMP column
in microseconds, from 25 September 2026:
SELECT timestampFROM tradesWHERE timestamp BETWEEN '2026-09-25T09:56:06.681999'AND '2026-09-25T09:56:06.687000'LIMIT 18;
Trades arrive in bursts: several fills share the same microsecond, and then the feed moves on a millisecond at a time.
Half of these rows cost a single bit, because a burst of fills shares one microsecond and the gap stops changing. Nine rows at 1 bit, two at 9 and five at 16, plus the two full timestamps, come to 235 bits against 1,152 for the plain form: a 4.9x saving, on tick data rather than a metric sampled on a timer.
Those eighteen rows are the best case for this feed, because they sit inside a single burst. Between bursts the feed can go quiet for a few milliseconds, and any gap that differs from the one before it by more than about two milliseconds needs the 36-bit bucket. A longer run includes plenty of those. The last thousand trades on 28 September 2026 took 25,605 bits against 64,000, about 2.5x. The script at the end of this post repeats that calculation on whatever the table holds when you run it.
Regular series: one bit per timestamp, mostly
Sampled metrics, downsampled series and anything produced on a schedule look nothing like the ones above. They're meant to have the same gap every time, so the delta of delta is zero and each timestamp costs one bit. This is the case most dashboards, sensors and sampled queries fall into, and it's where Gorilla does best.
In practice the gap is only mostly the same. This query downsamples forty-one
seconds of the same trades table to one row per second, which is what a
dashboard panel would ask for. When a QWP client runs it, the result travels
back over QWP, so this time it's the server doing the encoding:
SELECT timestamp, count()FROM tradesWHERE timestamp BETWEEN '2026-09-28T11:02:03'AND '2026-09-28T11:02:43.999999'SAMPLE BY 1s;
It returns forty rows, not forty-one. Nobody traded during 11:02:37, and without
a FILL clause SAMPLE BY doesn't emit a row for an empty interval. Here are
the rows around the missing second:
| Timestamp | Gap (µs) | Delta of delta (µs) | Bits |
|---|---|---|---|
| 11:02:35 | 1,000,000 | 0 | 1 |
| 11:02:36 | 1,000,000 | 0 | 1 |
| 11:02:38 | 2,000,000 | +1,000,000 | 36 |
| 11:02:39 | 1,000,000 | -1,000,000 | 36 |
| 11:02:40 | 1,000,000 | 0 | 1 |
One gap is twice as long as the rest, so the delta of delta jumps going into it and jumps back coming out, and both values need the 36-bit bucket. Every other row costs one bit.
| 40 rows | One second missing | No second missing |
|---|---|---|
| First two, in full | 128 bits | 128 bits |
| Zero DoD, 1 bit each | 36 | 38 |
| Widest bucket, 36 bits each | 2 | 0 |
| Gorilla total | 236 bits | 166 bits |
| Plain total | 2,560 bits | 2,560 bits |
| Ratio | 10.8x | 15.4x |
One quiet second costs 70 bits, nearly twice the 36 bits of the thirty-six unchanged rows put together, and the column still takes 91% fewer bits than the plain form.
It's 15x rather than the 64x that one bit per row would suggest, because of the two timestamps sent in full. Those 128 bits are fixed per column and per batch, so they matter over forty rows and fade over a thousand: a thousand timestamps with no missing interval cost 1,126 bits against 64,000, about 57x.
So the same encoding gives 2.5x on a bursty trade feed and up to 57x on a regular series, and it never makes a column bigger. The one thing that stops it is a value it can't fit at all. On ingestion the client encodes and the server decodes, and for query results it's the other way round.
When Gorilla falls back to plain 64-bit timestamps
The widest bucket holds the delta of delta in a 32-bit signed integer, so it covers roughly plus or minus 2.1 billion. What that means in wall clock time depends entirely on the unit of your timestamp column:
| Column type | 32-bit range, in time |
|---|---|
TIMESTAMP, microseconds | about 35 minutes |
TIMESTAMP_NS, nanoseconds | about 2.1 seconds |
At microsecond resolution it takes a pause of more than about 35 minutes inside a single batch, a market close for example, to go past that ceiling.
With nanosecond timestamps, the same jitter is a thousand times larger as a
number, so a change that fits a small bucket in microseconds needs the widest
one in nanoseconds. Eighteen consecutive rows of fx_trades, the demo's FX feed
at nanosecond resolution, put every delta of delta in the 36-bit bucket: 704
bits against 1,152, about 1.6x, where the microsecond trades burst earlier got
4.9x. The ceiling is also a thousand times closer, and real data reaches it.
Here are four consecutive rows we pulled from fx_trades, where the feed pauses
for a little over two seconds:
| Timestamp | Gap (ns) | Delta of delta (ns) | Bucket |
|---|---|---|---|
| 14:04:09.476727131 | in full | ||
| 14:04:09.486301108 | 9,573,977 | in full | |
| 14:04:09.486640354 | 339,246 | -9,234,731 | 36 bits |
| 14:04:11.645017351 | 2,158,376,997 | 2,158,037,751 | doesn't fit |
The third row already needs the widest bucket, but a change of 9 milliseconds is well inside it. The fourth is a pause of 2.158 seconds, a delta of delta of 2,158,037,751 nanoseconds against a ceiling of 2,147,483,647. It misses by about 10.6 milliseconds, and there's no bucket left to put it in.
When that happens, the encoder drops Gorilla for that timestamp column in that batch and sends every value as a plain 64-bit integer. The decision is per column and per batch, so the next batch can use Gorilla again if its timestamps allow it.
The encoder also needs at least three non-null values, since with two there's no delta of delta to encode. Beyond that, a column that fits can't lose: the widest bucket costs 36 bits against 64, so as long as every value fits, the Gorilla form is smaller than the plain one. The QWP ingress specification has the byte-level layout, and the egress specification covers query results.
Run the Gorilla calculation on your own tables
All of the above is reproducible against our public demo instance. This pulls the last thousand timestamps from a table over the REST API and works out the buckets:
import json, urllib.parse, urllib.requestfrom datetime import datetimeTABLE = "trades" # or fx_trades, for nanosecond timestampsSQL = f"SELECT timestamp FROM {TABLE} LIMIT -1000"URL = ("https://demo.questdb.io/api/v1/sql/execute?query="+ urllib.parse.quote(SQL))# the demo rejects urllib's default User-Agent, so send our ownreq = urllib.request.Request(URL, headers={"User-Agent": "gorilla-check"})rows = json.load(urllib.request.urlopen(req))["dataset"]EPOCH = datetime(1970, 1, 1)def to_int(s):# "2026-09-25T09:56:06.681999Z" as an integer in the column's own unit:# microseconds for TIMESTAMP, nanoseconds for TIMESTAMP_NSwhole, frac = s.rstrip("Z").split(".")secs = int((datetime.fromisoformat(whole) - EPOCH).total_seconds())return secs * 10 ** len(frac) + int(frac)def cost(dod):if dod == 0: return 1if -64 <= dod <= 63: return 9if -256 <= dod <= 255: return 12if -2048 <= dod <= 2047: return 16if -2**31 <= dod <= 2**31 - 1: return 36return None # doesn't fit: the whole column falls back to plain 64-bitts = [to_int(r[0]) for r in rows]bits = 128 # the first two timestamps go in fullfor i in range(2, len(ts)):dod = (ts[i] - ts[i-1]) - (ts[i-1] - ts[i-2])c = cost(dod)if c is None:print("fallback to plain 64-bit at row", i)bits = 64 * len(ts)breakbits += cprint(bits, "bits against", 64 * len(ts), "plain")
Change TABLE to another table on the demo instance, or point the script at
your own QuestDB, to see how your own timestamp columns would encode.
Next: varints, ZigZag and bit packing
Gorilla works because a timestamp column usually moves forward in steps the encoder can predict. Most other columns make no such promise, so the rest of a QWP frame uses encodings that assume less about the data. Varints store a small integer in one byte instead of eight, and that's how symbol IDs, row counts and the lengths of table and column names travel. ZigZag encoding keeps a small negative number small enough for a varint. Bit packing fits eight booleans into a byte. The encoding primitives section of the spec has the byte-level rules, and part II of this series, coming soon, shows how each one works and what it saves.