unlisted · measured · reproducible
pgvector benchmarks: the real numbers
Every figure below was measured with our open bench.py harness on real Rivestack-class nodes, over the production client path (TLS, a same-datacenter client), on 20 July 2026. Nothing here is modeled or rounded up. Each QPS number carries its recall@10 and its client count, because on approximate nearest-neighbour search those never travel alone: peak throughput and minimum latency are opposite ends of the same curve.
Two things we won't hide
1. cpx is shared vCPU, not dedicated. Every region now runs the AMD cpx line, fast per core and reliably in stock, but the vCPUs are shared (AMD EPYC), not isolated. Only Hetzner's ccx line is dedicated, and we don't use it. The win over the older Intel cx line is faster cores plus availability, not CPU isolation. (US and Asia run the same cpx family at slightly different sizes.)
2. The client was never the bottleneck. Every load generator has ≥ the DB's core count, and throughput scaled with concurrency on every tier (e.g. Scale roughly doubled from 2,098 → 4,031 QPS going 4 → 8 clients on its 8 cores). If the client were saturating, QPS would plateau. It didn't.
Solo
1 vCPU / 2 GB · cpx12 (AMD · shared vCPU) · $29/mo250,000 × 1536-dim vectors · HNSW m=16 ef_construction=64 cosine · index build 889s · load generator cpx32 (4 vCPU)
| ef_search | recall@10 | 1 client QPS · p50 | 4 clients QPS · p50 | 8 clients QPS · p50 | 16 clients QPS · p50 |
|---|---|---|---|---|---|
| 10 | 0.738 | 972 · 1ms | 1446 · 2.7ms | 1347 · 5.5ms | 1172 · 12.5ms |
| 20 | 0.788 | 847 · 1.1ms | 1261 · 3ms | 1263 · 5.8ms | 1285 · 11.5ms |
| 40 | 0.838 | 825 · 1.1ms | 1166 · 3.3ms | 1250 · 5.9ms | 1458 · 9.7ms |
| 80 | 0.925 | 712 · 1.3ms | 983 · 3.8ms | 886 · 8.2ms | 917 · 15.8ms |
| 120 | 0.964 | 636 · 1.4ms | 832 · 4.4ms | 781 · 9.7ms | 689 · 20.9ms |
| 200 | 0.974 | 53 · 13.4ms | 163 · 18ms | 222 · 27.1ms | 345 · 35ms |
Highlighted row ≈ recall@10 0.90 to 0.95, the usual production operating band.
Starter
2 vCPU / 4 GB · cpx22 (AMD · shared vCPU) · $49/node250,000 × 1536-dim vectors · HNSW m=16 ef_construction=64 cosine · index build 293s · load generator cpx32 (4 vCPU)
| ef_search | recall@10 | 1 client QPS · p50 | 4 clients QPS · p50 | 8 clients QPS · p50 | 16 clients QPS · p50 |
|---|---|---|---|---|---|
| 10 | 0.753 | 392 · 2.4ms | 1463 · 2.6ms | 1932 · 4ms | 2088 · 7.1ms |
| 20 | 0.813 | 398 · 2.3ms | 1397 · 2.7ms | 1882 · 4.1ms | 1996 · 7.6ms |
| 40 | 0.881 | 342 · 2.9ms | 1291 · 3ms | 1638 · 4.7ms | 1770 · 8.4ms |
| 80 | 0.930 | 369 · 2.5ms | 1185 · 3.2ms | 1497 · 5ms | 1599 · 9.3ms |
| 120 | 0.968 | 356 · 2.6ms | 1115 · 3.4ms | 1356 · 5.6ms | 1360 · 10.7ms |
| 200 | 0.988 | 214 · 4.5ms | 717 · 5.4ms | 838 · 9ms | 936 · 15.6ms |
Highlighted row ≈ recall@10 0.90 to 0.95, the usual production operating band.
Growth
4 vCPU / 8 GB · cpx32 (AMD · shared vCPU) · $85/node250,000 × 1536-dim vectors · HNSW m=16 ef_construction=64 cosine · index build 122s · load generator cpx62 (16 vCPU)
| ef_search | recall@10 | 1 client QPS · p50 | 4 clients QPS · p50 | 8 clients QPS · p50 | 16 clients QPS · p50 |
|---|---|---|---|---|---|
| 10 | 0.818 | 604 · 1.6ms | 2278 · 1.7ms | 3326 · 2.3ms | 3961 · 3.7ms |
| 20 | 0.854 | 602 · 1.6ms | 2291 · 1.7ms | 3101 · 2.4ms | 3712 · 4ms |
| 40 | 0.903 | 530 · 1.8ms | 2126 · 1.8ms | 2907 · 2.6ms | 3370 · 4.4ms |
| 80 | 0.939 | 505 · 1.9ms | 1989 · 1.9ms | 2675 · 2.8ms | 2949 · 5.1ms |
| 120 | 0.949 | 479 · 2ms | 1812 · 2.1ms | 2340 · 3.2ms | 2574 · 5.8ms |
| 200 | 0.968 | 316 · 3ms | 1238 · 3.1ms | 1468 · 5.2ms | 1497 · 9.8ms |
Highlighted row ≈ recall@10 0.90 to 0.95, the usual production operating band.
Scale
8 vCPU / 16 GB · cpx42 (AMD · shared vCPU) · $159/node250,000 × 1536-dim vectors · HNSW m=16 ef_construction=64 cosine · index build 76s · load generator cpx62 (16 vCPU)
| ef_search | recall@10 | 1 client QPS · p50 | 4 clients QPS · p50 | 8 clients QPS · p50 | 16 clients QPS · p50 |
|---|---|---|---|---|---|
| 10 | 0.803 | 522 · 1.8ms | 2098 · 1.9ms | 4031 · 1.9ms | 5815 · 2.6ms |
| 20 | 0.838 | 510 · 1.9ms | 1985 · 1.9ms | 3928 · 2ms | 5539 · 2.7ms |
| 40 | 0.898 | 461 · 2.1ms | 1847 · 2ms | 3600 · 2.1ms | 5091 · 3ms |
| 80 | 0.953 | 455 · 2.1ms | 1724 · 2.2ms | 3323 · 2.3ms | 4465 · 3.4ms |
| 120 | 0.973 | 398 · 2.4ms | 1534 · 2.4ms | 3058 · 2.5ms | 3883 · 3.9ms |
| 200 | 0.983 | 266 · 3.6ms | 1076 · 3.6ms | 2068 · 3.7ms | 2418 · 6.3ms |
Highlighted row ≈ recall@10 0.90 to 0.95, the usual production operating band.
Scale: 1,000,000 × 1536 (large index)
cpx42 · 8 vCPU / 16 GB · ~6 GB HNSW index built in 544s (9 min) and held hot in RAM. At bigger N the same ef_search gives lower recall (more true neighbours to find), so read recall per row.
| ef_search | recall@10 | 1 client | 4 clients | 8 clients | 16 clients |
|---|---|---|---|---|---|
| 10 | 0.561 | 479 · 2ms | 1920 · 2ms | 3765 · 2ms | 5446 · 2.8ms |
| 20 | 0.623 | 403 · 2.4ms | 1866 · 2.1ms | 3554 · 2.1ms | 4921 · 3.1ms |
| 40 | 0.674 | 431 · 2.2ms | 1682 · 2.3ms | 3321 · 2.3ms | 4294 · 3.5ms |
| 80 | 0.738 | 386 · 2.4ms | 1493 · 2.5ms | 2936 · 2.5ms | 3600 · 4.2ms |
| 120 | 0.778 | 354 · 2.6ms | 1397 · 2.6ms | 2664 · 2.7ms | 3176 · 4.7ms |
| 200 | 0.845 | 273 · 3.4ms | 1087 · 3.4ms | 2095 · 3.5ms | 2410 · 6.2ms |
# re-validated 10 August 2026
In August we changed the fleet's connection and memory defaults: every dedicated node now runs PgBouncer with 20 to 40 transaction slots per database (500 to 10,000 pooled client connections by tier) and a tighter work_mem of 8 to 32 MB. Because those settings could plausibly move the numbers above, we re-provisioned all four tiers from scratch and re-measured with the same harness. Like-for-like operating points came back within shared-vCPU run variance, so the July tables stand. Two additional findings, both measured:
- - work_mem A/B, zero difference. Filtered vector queries with
hnsw.iterative_scan = relaxed_orderat 10% and 1% selectivity, measured order-interleaved after warmup at both the new and the old work_mem on every tier: recall@10 exactly equal, p50 and p95 within noise. The scan-memory bound is never reached at these dataset sizes. - - Past the pool, queries queue instead of failing. We ran double each tier's transaction slots (80 clients on Scale's 40). Zero connection errors; backends pin at the cap and the extra clients wait, trading p50 for admission.
New measurements the July run didn't cover (pgvector 0.8.6):
| tier | dataset | ef_search | recall@10 | clients | QPS · p50 |
|---|---|---|---|---|---|
| Solo | 100k × 1536 | 40 | 0.959 | 1 | 489 · 2ms |
| Solo | 100k × 1536 | 20 | 0.849 | 20 | 1,417 · 13.7ms |
| Starter | 250k × 1536 | 80 | 0.927 | 25 | 1,468 · 16.8ms |
| Growth | 500k × 1536 | 80 | 0.834 | 30 | 2,493 · 11.5ms |
| Growth | 500k × 1536 | 200 | 0.934 | 30 | 1,768 · 16.3ms |
| Scale | 1M × 1536 | 80 | 0.714 | 40 | 3,455 · 11.1ms |
| Scale | 1M × 1536 | 200 | 0.829 | 40 | 2,125 · 17.8ms |
| Scale | 1M × 1536 | 10 | 0.558 | 80 | 4,629 · 16.7ms2× the pool: every client past 40 slots queues; nothing errors |
A 500k × 1536 index (~3 GB) builds and serves hot on Growth (8 GB); Scale's 1M row above is the same large-index claim as July, now also measured at 40 and 80 clients. Read recall per row, as always: the 4,629 QPS row is a low-recall operating point shown for the queueing behavior, not a headline.
# methodology
- - Data: Gaussian-mixture clusters (structure dominates, not uniform random noise). Queries are small perturbations of real points, giving a well-defined recall@10 against exact brute-force KNN.
- - Index: pgvector 0.8.5 HNSW, m=16, ef_construction=64, cosine. PostgreSQL 17.10.
- - Tuning: each node uses Rivestack's per-size defaults: shared_buffers 25% RAM, effective_cache_size 75%, parallel workers = vCPU, random_page_cost 1.1. The July tables ran work_mem = RAM/64; since August the default is 8 to 32 MB by tier, and the re-validation above measured no difference between the two.
- - ef_search is set with
SET LOCALinside each query's transaction (a plain session SET is silently dropped by PgBouncer transaction pooling). - - Client path: a separate same-datacenter VM (≥ the DB's core count) over TLS. Measured network floor 0.34 to 0.51 ms per round-trip. Throughput uses one process per client (no GIL).
- - What we don't claim: peak QPS and minimum p50 at the same time; competitor numbers we didn't run; a single number for all regions.
# reproduce it
On a same-region VM with Python, numpy and psycopg2:
# harness: https://github.com/Rivestack/pgvector-bench
export DSN='postgres://USER:PASS@YOUR-DB.eu.db.rivestack.io:6432/db?sslmode=verify-full'
export N=250000 EFS='10,20,40,80,120,200' CONCS='1,4,8,16' DUR=12
python3 bench.py
# prints: ef | recall@10 | per-concurrency QPS/p50/p95/p99Point it at your own database and dataset, and the numbers you get are the numbers you get. That's the whole idea.
Measured & signed by Yasser, Rivestack, 20 July 2026. Re-validated under the current fleet configuration on 10 August 2026.
Runs on the AMD cpx line, which every region now uses (cpx is shared vCPU, not dedicated, see above). Re-run anytime with the harness; if your numbers differ, tell us and we'll publish the correction.