⏱️ Reading time: 16 min

Turbopuffer, the serverless vector database used by Cursor and Notion, published an article on September 30, 2026 titled “RIP, vector database” announcing that its internal architecture no longer revolves around the vector index.

📑 En este artículo
  1. TL;DR
  2. What is a serverless vector database?
  3. Why it matters
  4. How it works under the hood
    1. Object storage as the source of truth
    2. Tiered cache: NVMe and memory
    3. How vectors are indexed: SPANN and SPFresh
  5. From v1 to v3: the evolution of turbopuffer
    1. v1: just an ID and a vector
    2. v2: attribute filters and full-text search
    3. The problem with having a vector index as primary
    4. v3: the ANN index becomes secondary
  6. Practical examples and how to get started
  7. Real-world use cases
  8. Common mistakes and best practices
  9. Comparison with alternatives
  10. Going deeper
  11. Frequently Asked Questions
    1. What is a serverless vector database and how is it different from pgvector?
    2. Why doesn’t turbopuffer use an HNSW index like most vector engines?
    3. What exactly does turbopuffer v3 change compared to v2?
    4. Is a serverless vector search engine useful for something other than search?
    5. What happens to turbopuffer’s latency when a namespace hasn’t been used in a while?
  12. References

The company is migrating to turbopuffer v3, a rewrite that turns the ANN index (the one that handles similarity search) into “just one more” of several secondary indexes, instead of the system’s central axis. It’s the perfect case for understanding, end to end, how the storage of a modern vector database is designed.

TL;DR

  • On September 30, 2026, turbopuffer announced that the ANN index is no longer the axis of its storage.
  • Object storage is the source of truth; a tiered NVMe and memory cache answers queries quickly.
  • SPANN and SPFresh group vectors into a hierarchical cluster tree instead of a graph like HNSW.
  • The vector-primary design held back queries like GROUP BY and aggregations over the same data.
  • Cursor, Notion, and Linear’s sync engine already run on this serverless architecture.

What is a serverless vector database?

A vector database is a system that stores numerical representations (embeddings) of text, images, or audio and answers queries by similarity, not exact equality. The serverless variant stores that data in cheap object storage and allocates compute on demand, with no fixed servers to maintain or administer.

Cursor and Notion have run on turbopuffer since its launch as a serverless vector database. Foto de Creative Minds Factory en Unsplash

Why it matters

For years, setting up a vector database meant standing up your own cluster: sizing memory for the HNSW index, replicating it, and paying for reserved capacity even when traffic was intermittent. The serverless model flips that equation: the data stays in cheap object storage, the same kind of storage S3 uses, and compute is allocated only when a query arrives.

That separation between storage and compute is the same one adopted by modern data warehouses like Snowflake or BigQuery, now applied to vectors. The result is that a nearly idle namespace costs almost nothing, because storing a vector in object storage is orders of magnitude cheaper than keeping it loaded in the memory of a dedicated cluster.

This matters especially for RAG (retrieval-augmented generation) applications and AI agents, where each user, document, or conversation may need its own isolated namespace. Keeping thousands of small namespaces alive in memory at all times isn’t economically viable; serving them on demand from object storage is.

How it works under the hood

Object storage as the source of truth

In turbopuffer, each namespace (the logical unit where a customer’s documents live) is stored as a set of immutable files in object storage. When the client writes new vectors, the system doesn’t update them in place: it creates a new file and appends it to the namespace, similar to how Git appends commits instead of rewriting history.

This immutability is what makes the serverless cost model possible. Object storage doesn’t charge for keeping a server running, only for the gigabytes stored and for read and write operations. But there’s a tradeoff: the latency of a direct read from object storage is measured in tens or hundreds of milliseconds, too slow to serve an interactive search.

Tiered cache: NVMe and memory

To solve the latency problem, turbopuffer adds two cache tiers between the client and object storage: a RAM memory tier, the fastest but also the most expensive and limited, and an NVMe SSD tier, slower than RAM but much cheaper and with enough capacity to store entire namespaces that don’t fit in memory.

When a query arrives, the engine first looks for the data in memory. If it’s not there, a “cache miss”, it checks NVMe. If it’s not there either, only then does it read from object storage and, in the process, promotes that data to the upper cache tiers so the next query is faster. It’s the same principle a CPU uses with its L1, L2, and L3 caches, applied to a distributed system.

flowchart TD
A["Client"] --> B["turbopuffer query engine"]
B --> C["Memory cache"]
C --> D["NVMe cache"]
D --> E[("Object storage: source of truth")]
subgraph "Cache tiers"
C
D
end
sequenceDiagram
participant C as Client
participant Q as Query engine
participant M as Memory cache
participant N as NVMe cache
participant O as Object storage
C->>Q: query top_k=10
Q->>M: checks hot cache
alt data in memory
M-->>Q: returns result
else cache miss
Q->>N: checks NVMe
alt data in NVMe
N-->>Q: returns result
else total miss
Q->>O: reads from object storage
O-->>Q: returns data
end
end
Q-->>C: returns final result

How vectors are indexed: SPANN and SPFresh

For similarity search, turbopuffer doesn’t use a graph-based index like HNSW, the kind most self-hosted vector databases use. It uses a hierarchical clustering index, based initially on SPANN and later migrated to SPFresh to support incremental indexing without rebuilding the whole tree on every write.

flowchart TD
R["Root centroid"] --> L1["Leaf centroid 1"]
R --> L2["Leaf centroid 2"]
R --> L3["Leaf centroid 3"]
L1 --> V1["Vector C0L0"]
L1 --> V2["Vector C0L1"]
L2 --> V3["Vector C1L0"]
L2 --> V4["Vector C1L1"]
L3 --> V5["Vector C2L0"]
L3 --> V6["Vector C2L1"]

The idea is to group vectors into clusters, group the centroids of those clusters into higher-level clusters, and repeat the process until reaching a single root. An HNSW-type graph is hard to split into independent files because any node can connect to any other; a cluster tree, on the other hand, lends itself to living in separate files in object storage, because each branch is relatively independent of the others.

Each vector is identified by what turbopuffer calls an “ANN address”: a ClusterId that indicates which group it belongs to, and a LocalId that identifies it within that group. In turbopuffer’s first two versions, that ANN address was literally the primary key of the entire system: the other indexes, attribute filters, text search, pointed to ANN addresses, not document IDs.

From v1 to v3: the evolution of turbopuffer

v1: just an ID and a vector

turbopuffer started out as a pure serverless vector database: each document was just an ID and a vector. The entire architecture, from the file format to the query plan, revolved around the SPANN/SPFresh clustering tree described above. That specialization is what let it offer low prices: by not carrying the generality of a relational database, it could optimize every byte stored in object storage specifically for similarity search.

Customer demand forced the first expansion: being able to filter a vector search by attributes, for example “vectors similar to this query, but only from user X”. turbopuffer solved this with a classic inverted index, the same type of structure text search engines use: an attribute-value key points to the list of ANN addresses that contain it.

Full-text search with BM25 ranking arrived through the same logic: for each term, the index stores the list of documents that contain it along with the term’s frequency and the document’s length, the two pieces of data BM25 needs to score relevance. From there, turbopuffer added aggregations, regex search, fuzzy matching, and sparse vectors, all built on the same layout with the ANN index as primary.

The problem with having a vector index as primary

turbopuffer itself admits that this design worked very well for what it was meant to do: fast, cheap similarity search. But over time it became a straitjacket for everything else. Queries like GROUP BY or attribute aggregations, common in SQL, fit poorly when the underlying storage structure is optimized for nearest neighbors, not for grouping rows by a column.

The reason is structural: if every secondary index points to an ANN address, any operation that doesn’t first pass through the clustering tree has to indirectly reconstruct the relationship between documents. It’s the equivalent of always having to check a book’s index to find a page, even when you already know the chapter number.

v3: the ANN index becomes secondary

The solution turbopuffer started implementing on September 30, 2026 flips the hierarchy: the company is migrating toward a more flexible primary index, separate from the vector clustering tree, and turns ANN into just another secondary index, on the same level as the attribute filter or text search.

📌 Note: turbopuffer is still in the middle of that migration and is documenting it publicly, article by article, so the final design of v3’s primary index isn’t settled yet.

What the announcement does confirm is the goal: that the same storage serve both vector search and SQL-type queries better, without one taking efficiency away from the other.

SPANN and SPFresh organize vectors into cluster trees, not HNSW-type graphs. Foto de Ilya Pavlov en Unsplash

Practical examples and how to get started

turbopuffer is a managed service, not a package you install locally, but it does have official clients in Python, TypeScript, and Go that work the same way on Windows, macOS, and Linux because they talk to the API over HTTPS. To follow these examples you need an account and an API key from turbopuffer.com.

pip install turbopuffer

This command installs the official Python client and works the same way on Windows, macOS, or Linux because it’s a pure package over HTTPS, with no native dependencies to compile.

import os
import turbopuffer as tpuf

tpuf.api_key = os.environ["TURBOPUFFER_API_KEY"]

ns = tpuf.Namespace("demo-documents")
ns.upsert(
    ids=[1, 2],
    vectors=[[0.12, 0.98, 0.33], [0.45, 0.21, 0.88]],
    attributes={
        "titulo": ["Introduction to RAG", "Semantic search"],
        "idioma": ["es", "es"],
    },
)

This block creates, or reuses, the demo-documents namespace and uploads two 3-dimensional vectors with their attributes. In a real case the vectors would come from an embeddings model and would have hundreds of dimensions, not three.

ns.query(
    vector=[0.10, 0.95, 0.30],
    top_k=2,
    filters=("idioma", "Eq", "es"),
)

This call looks for the 2 vectors closest to the query vector, but only among documents in Spanish. It returns a list of results sorted by distance. The expected output, summarized and with field names that may vary depending on the client version, looks something like this:

[
  {"id": 2, "dist": 0.04, "attributes": {"titulo": "Semantic search"}},
  {"id": 1, "dist": 0.11, "attributes": {"titulo": "Introduction to RAG"}}
]

To confirm the namespace exists and how many vectors it has, turbopuffer exposes a web dashboard at turbopuffer.com in addition to a metadata endpoint. There’s no direct way to paste the exact response of that endpoint here because the format changes between client versions; the most reliable approach is to check the official documentation before automating a check against specific fields.

💡 Tip: if your application creates a namespace per user or per conversation, a common pattern in RAG, the serverless model is the difference between paying for thousands of idle clusters and paying almost nothing for the namespaces nobody queries this week.

Real-world use cases

Cursor, the AI code editor, uses turbopuffer to index and search entire codebases in real time. Notion uses it for semantic search across its users’ documents. Neither is a lab case: both run in production with real traffic and namespaces that grow every day.

The most interesting case for understanding where turbopuffer v3 is headed is Linear, which uses the system for something that isn’t search: its real-time sync engine between client and server. That “non-search” use is exactly what the vector-primary design of v1 and v2 didn’t facilitate, and it’s part of the motivation behind the redesign toward a flexible primary index.

Common mistakes and best practices

  • Treating it as a general-purpose relational database, it’s optimized for object storage and similarity queries; forcing complex joins or multi-table transactions isn’t its strength.
  • Ignoring cache warm-up, a namespace that hasn’t been queried in a while lives only in object storage; the first query after a period of inactivity will be slower because it has to promote the data to NVMe and memory before it can respond quickly.
  • Filtering by high-cardinality attributes without measuring the impact, the inverted index works best with a manageable number of distinct values; an attribute like user_id with millions of unique values generates very fragmented posting lists.
  • Underestimating the cost of object storage reads, storage itself is cheap, but each read operation has its own cost; a query pattern that ignores the cache and always hits the cold tier can end up more expensive than expected.
  • Not separating namespaces per customer when the use case calls for it, namespace isolation is what allows scaling to thousands of customers without one affecting another’s latency; mixing everything into one giant namespace reduces that advantage.

Comparison with alternatives

Not every application needs a dedicated serverless vector search engine. The following table compares the most common options in 2026.

OptionWhen to use itAdvantageLimitation
Serverless on object storage (turbopuffer)Many namespaces with intermittent traffic, multi-tenant RAG, agentsNear-zero cost at rest, scales to thousands of namespacesColder latency than an always-in-memory index
pgvector (PostgreSQL extension)Relational data already exists and vector volume is moderateA single engine for SQL and vectors, no new system to operateThe index lives inside the Postgres cluster itself: scaling a lot means scaling everything
Self-hosted engine with HNSW (e.g. Qdrant, Weaviate)Constant low latency and full infrastructure controlAlways warm, no cache warm-upYou have to size and pay for the cluster even when traffic is low
FAISS embedded in the applicationPrototypes and small datasets that fit in a single process’s memoryZero external infrastructure, full control within the process itselfNot distributed or persistent on its own

Going deeper

The decision to use hierarchical clustering, SPANN and SPFresh, instead of an HNSW graph isn’t just a technical preference: it’s what makes it viable to store the index in object storage. An HNSW graph needs to jump from node to node following pointers that can be anywhere in the file; that’s fast in memory but extremely expensive if every jump means a read from a system with hundreds of milliseconds of latency. A cluster tree, by contrast, is traversed level by level, and each level can be stored as a contiguous range of sorted keys.

That also explains why turbopuffer models its storage as a sorted, unique key-value map, with keys composed of ClusterId and LocalId. A sorted key enables efficient range scans: asking for “all the vectors in cluster 7” is, in storage terms, asking for a contiguous range of keys, not a scattered search.

The shift toward v3 generalizes that same sorted-key idea, but without assuming the ANN has to be the root of that hierarchy. If the primary key becomes, for example, the document’s ID instead of its ANN address, grouping by attribute or running a GROUP BY no longer requires going through the vector clustering tree first. The ANN index becomes a secondary index that points to that primary key, exactly the reverse of how it worked in v1 and v2.

Your next step: create a free account at turbopuffer.com, upload 50 vectors from a test dataset using the Python client, and run a query with an attribute filter to see in practice the difference between the ANN index and the inverted index.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What is a serverless vector database and how is it different from pgvector?

Both solve similarity search, but a serverless vector database like turbopuffer stores data in object storage and scales compute per query, while pgvector runs inside a PostgreSQL instance that has to be sized and kept running all the time.

Why doesn’t turbopuffer use an HNSW index like most vector engines?

Because an HNSW graph doesn’t lend itself well to living fragmented across object storage files: every jump between nodes can land anywhere in the index. The SPANN and SPFresh clustering tree is traversed level by level and stored as sorted ranges, something much friendlier to object storage.

What exactly does turbopuffer v3 change compared to v2?

v2 used the ANN address, ClusterId and LocalId, as the primary key for the entire system, and the other indexes pointed to that address. v3 separates that relationship: the ANN index becomes secondary, like the attribute index or the full-text index, instead of being the axis of all storage.

Yes. Linear uses it as a real-time sync engine between client and server, not for search. That’s exactly the kind of use case the v3 redesign aims to make easier, by no longer forcing everything through the vector index.

What happens to turbopuffer’s latency when a namespace hasn’t been used in a while?

If the data lives only in object storage because no cache tier has it loaded, the first query will take longer while the system promotes that data to NVMe and memory. Subsequent queries on the same namespace are faster because the data is already cached.

References

📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day. @programacion

Featured image: Foto de imgix en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: ProgrammingTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.