Databricks Releases Lakebase Search for Postgres
Databricks has launched Lakebase Search on AWS and Azure, introducing native vector and text search extensions to Postgres to eliminate complex ETL pipelines for AI applications.

Databricks has announced the general availability of Lakebase Search on AWS and Azure, integrating high-performance search capabilities directly into Lakebase Postgres. The release features two new extensions: lakebase_vector for approximate nearest neighbor search and lakebase_text for BM25 full-text search. By combining semantic, keyword, and hybrid search within the operational database, the update removes the need for developers to maintain separate search engines and complex ETL pipelines.
The system is designed to address the scaling limitations of traditional tools like pgvector. On the VectorDBBench LAION 100M-vector benchmark, lakebase_vector delivered twice the throughput of the next-best system at a cost four times lower than cloud Postgres running pgvector. It achieved a 97 percent recall rate with a P99 latency of 71 milliseconds. The architecture achieves this by decoupling storage from compute, utilizing hierarchical IVF clustering and binary quantization via RaBitQ to compress 768-dimensional float32 vectors down to roughly one bit per dimension, which is about 32 times smaller.
For practitioners, the serverless architecture scales to zero, meaning users pay for query activity rather than idle data volume. Cold starts take approximately 1.13 seconds at the P90 interval for a first query on 100 million vectors, and an entire 100-million-vector dataset can be served on just a single Lakebase Compute Unit. Additionally, index builds are offloaded from the primary database to distributed engines like Spark, preventing performance degradation during updates.
The update also introduces lakebase_text, which scores terms using global inverse document frequency to outperform standard Postgres tsvector and GIN indexes. Early adopters like Conexiom have used these hybrid capabilities to run searches across more than 100 million rows. Conexiom reported cutting its database infrastructure spend threefold while achieving five times the throughput of its previous pgvector setup.
This is our own summary of reporting by Databricks AI



