Accelerating LLM Inference with Prompt Caching for Open‑Source Models on Databricks
Summary
Databricks now supports prompt caching for open-source models across all workloads, automatically accelerating LLM inference by reusing repeated prompt prefixes. This feature boosts throughput by 2.5x and reduces P50 latency by 3x for models like GPT-OSS, with no setup required.
Summary generated by brickster.ai. For the full article, follow the source link above.
More from Databricks Blog
NEAREST BY Join: Scaling Vector Search in Databricks Runtime
Databricks Runtime now features NEAREST BY, a new SQL join that executes exact or approximate batch vector searches directly against your Lakehouse data. Powered by a fused Photon operator with a custom blocked GEMM kernel, it turns standard liquid-clustered Delta tables into partition-pruned vector indexes with no external vector store to sync or manage.
Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs
The new REGISTER and UNREGISTER APIs provide a standard method to safely transfer open table format management between catalogs without copying data stored in customer-owned buckets. By requiring the original catalog to explicitly relinquish control before a transfer, this workflow prevents catalog lock-in and eliminates dangerous split-brain scenarios.
How Genie Ontology powers product development at Databricks
Databricks relies on Genie Ontology and Genie One to drive product development, combining human-curated semantics and learned knowledge across first- and third-party sources to analyze adoption and forecast trends. The platform extends these internal workflows with capabilities including custom skills, scheduled tasks, shareable agents, and MCP writes to external tools.
In Capital Markets, the Buy Side Runs on NAV. Finance Protects the Fee.
Databricks Genie One introduces an AI coworker for asset management finance that leverages backend agents to automate NAV reconciliation and fund data ingestion across a governed business ontology. Designed to protect net fee margins under strict regulatory standards like Form PF and GIPS, it provides finance teams with real-time, traceable morning NAV calculations and audit-ready proof for every valuation.
How to choose your first Genie Agents for maximum impact
Rank candidate Genie Agent workflows in minutes using a five-criteria rubric across impact, demand, data readiness, scope, and governance to ensure adoption scales rather than stalls. Field examples demonstrate how to apply this framework in practice to classify use cases into build now, shape it first, and hold off categories.
Real-Time Retail Intelligence: Building E-Commerce Recommendations with Lakebase and AI Search on Databricks
A complete reference architecture is now available for building a real-time, multi-stage e-commerce recommendation engine on Databricks using AI Search for candidate retrieval, Lakebase for online feature serving, and Model Serving for low-latency inference. By unifying batch precomputation and real-time scoring on a single governed platform, this design replaces fragmented ML stacks to eliminate glue code and deliver end-to-end lineage from raw clickstream to production predictions.
