> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.talqora.com/operate-at-scale/pricing-benchmark/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server. # Vector storage pricing benchmark This page explains how Talqora's public comparison is calculated. It is a planning tool for an early architecture decision, not a quote or a promise of a vendor's bill. ## Reference workload We normalize the comparison around a simple production-shaped workload: * 1,000,000 vectors at 1,536 dimensions * 1,000,000 retrieval queries per month * 1,000,000 vector writes per month * 4 KB of average metadata per vector * One region, no reserved capacity, and no enterprise contract The workload is intentionally explicit because vector platforms meter different things. A serverless service may meter read and write units, a managed cluster may meter CPU and memory, and a dimensions-based service may meter the total stored and queried dimensions. ## What the table means **Minimum / month** is the lowest published paid entry point we found for a comparable managed product. A `$0` value means a free tier or usage-only entry point, not unlimited production capacity. **Storage equivalent** is a readable conversion to dollars per GB. It does not replace the vendor's billing metric. For example, Cloudflare Vectorize publishes prices per stored and queried vector dimensions, while Weaviate publishes vector-dimension and GiB rates. The conversion uses float32 storage, where 1,536 dimensions represent 6,144 bytes before metadata and indexing overhead. **Billing model** describes the meter that actually controls the bill. Cluster-based products can have a low storage rate and still cost more because compute remains allocated while the workload is idle. ## Published inputs * [Pinecone pricing](https://www.pinecone.io/pricing/) publishes a $20 Builder plan and a $50 Standard monthly minimum, with usage meters for storage, read units, and write units. * [Weaviate pricing](https://weaviate.io/pricing) publishes a free tier and Flex from \$45/month, plus vector-dimension and storage rates. * [Cloudflare Vectorize pricing](https://developers.cloudflare.com/vectorize/platform/pricing/) publishes stored and queried vector-dimension rates, with usage included in Workers plans. * [Qdrant Cloud billing](https://qdrant.tech/documentation/cloud/pricing-payments/) explains resource usage units based on CPU, memory, and disk. The landing comparison uses a directional entry estimate, not a fabricated per-GB list price. * [Zilliz Cloud cost guidance](https://docs.zilliz.com/docs/understand-cost) distinguishes free, serverless, and dedicated modes and describes operation-based serverless billing. Turbopuffer and Talqora are represented as usage-oriented storage and query services. Their equivalent storage values are deliberately labeled estimates because neither service's complete workload price can be reduced to a single storage-only number. ## How to reproduce the estimate 1. Calculate stored dimensions: `vector_count × dimensions`. 2. Calculate queried dimensions: `monthly_queries × dimensions`. 3. Add metadata, replicas, backups, and egress only when the vendor bills them separately. 4. Apply the vendor's published minimum or free allowance. 5. Keep the result separate from embedding, reranking, OCR, and application compute costs. For Talqora, the same workload is modeled as dense storage plus retrieval usage. There is no idle read-node reservation in the baseline. Your actual bill depends on the selected plan, region, metadata volume, writes, queries, and egress. ## Why the numbers differ The comparison is not claiming that every backend has the same latency or feature set. It highlights the economic shape of each architecture: * **Usage-based services** follow storage and requests, which is useful for bursty workloads. * **Dimensions-based services** make vector volume easy to estimate but still meter query dimensions. * **Cluster services** reserve CPU, memory, replicas, and disk, which can be valuable for predictable latency but creates an idle baseline. * **Free tiers** are useful for development and demos, but production limits and support usually require a paid plan. Always verify the vendor's pricing page before purchasing. Regional rates, promotions, minimums, and included usage can change.