Skip to content
All topics

LLM

Recent items mentioning LLM across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

60 recent items1 release17 news32 videos10 community threads
What's happening in LLMAI synthesis · updated 1d ago

Databricks launched ai_decide, a new SQL and REST AI Function designed to deliver faster, lower-cost structured decisions on governed data than traditional LLMs 1. Across evaluation and agent workflows, practitioners are adopting tiered LLM judge patterns ahead of Jev's integration into MLflow 3.17 3 and deploying Unity Gateway tracing to eliminate token waste caused by ambiguous LLM inputs and broken agent tool calls 5.

Generated daily from the 5 most recent items mentioning LLM. Click any [N] to jump to the source.

Reddit

Databricks Micro Apps, App Spaces and Genie App Generator

Databricks launched App Spaces and Serverless Micro Apps at the end of last week (in Beta). Over the weekend, I tested migrating two of my production Databricks Apps to App Spaces. App Migration 1: Blocked by Zero Egress My Data Portfolio Project Creator required internet access to pull data stacks from live job postings. Then, the LLM needs internet access to research for open data sources to use. In standard Databricks Apps, this runs cleanly. In App Spaces, there is zero external internet egress including for LLMs. The app can reach internal workspace resources, but it cannot touch the outside web. If your app relies on third-party APIs, external databases, or web scraping, App Spaces is a non-starter until Databricks opens network egress. Workload 2: Success on Internal FinOps My client-facing DBU cost observability app reads workspace usage data and writes directly to Lakebase. Because it requires zero external network calls, the migration worked. Both the micro app compute and Lakebase scale to zero when idle. Cold starts take roughly 30 seconds (though in beta, you occasionally need a quick browser refresh once it spins up). For internal, low-frequency administrative tools, this turns a continuous monthly compute bill into pennies. (+ App Spaces, Genie App Generator, and Micro Apps are free while in beta) The next part is less about App Spaces and more just general best practice for Databricks Apps that I see people miss. Stop Using Delta Lake as an OLTP Database Databricks Apps are software applications, not batch analytics notebooks. If your app writes application state, session data, or row-level CRUD directly into analytical Delta tables, you need to rethink that design. For app transactions, use Lakebase (serverless Postgres which also scales to 0). Both are governed under UC, but Lakebase gives your app the low-latency transactional engine that application engineering actually requires. My Verdict on the Beta App Spaces solves the idle compute problem that has plagued Databricks Apps since launch. But until Databricks allows us to deploy directly from existing Git repos and opens external network egress, it remains limited to internal-only use cases. Curious on other peoples experience... How has it been for others? submitted by /u/OkImprovement7010 [link] [comments]

00OkImprovement70103d ago
Reddit

The deployment decalogue

I am an ML engineer, but I come from a software engineering background: years of full-stack work, with heavy DevOps and Terraform experience. I come from teams that deploy to production five times a day with real continuous deployment. And honestly? Pressing the button still feels weird sometimes. Every engineer knows that feeling, no matter how good the safety net is. So I wrote down the list that settles it. Ten commandments, one flow, written with data scientists and ML teams in mind, but it works for batch jobs, realtime inference, and LLMs alike. Answer honestly, and if all ten are true, you can ship to production anytime, in any form or way. submitted by /u/SuspiciousPavement [link] [comments]

00SuspiciousPavement3w ago
Reddit

How are you actually setting up AI/LLM evals in Databricks end-to-end? Looking for a step-by-step production workflow

We have a product where Databricks is our backend , and we’re now trying to properly evaluate and improve the quality of the AI-generated answers in our application. I’m looking for advice from people who have actually implemented LLM/GenAI evaluations in Databricks in production . I’d really appreciate an end-to-end, step-by-step explanation of how you would set this up from scratch. Specifically, how would you approach: Define what a “good answer” means Accuracy / correctness Relevance Completeness Groundedness / faithfulness to our data Hallucination rate Citation/source correctness Following user instructions Response consistency Latency and cost Create a benchmark / golden dataset Should we manually create a set of representative user questions? How many questions are enough to start? Should each question have an expected answer? Should we store expected SQL/results, expected sources, or just an expected natural-language response? Where should this benchmark dataset live in Databricks? How do you keep it updated as the product evolves? Set up automated evaluations What Databricks/MLflow tools should we be using today? MLflow evaluation? LLM-as-a-judge? Custom scorers? Human evaluation? How do you combine these rather than relying on one score? Evaluate RAG / data-grounded answers Our AI answers questions based on enterprise data in Databricks. How do you separately evaluate: Retrieval quality Whether the correct tables/documents were selected Context relevance Groundedness Final answer correctness Whether the model invented something that wasn’t in the retrieved data Evaluate text-to-SQL / analytics questions If a user asks something like: “What was average building occupancy last month?” should we evaluate: Generated SQL Tables selected SQL execution result Final natural-language answer separately? What is the recommended architecture for this? Guardrails Where should guardrails sit in the architecture? For example: Prevent hallucinated numbers Prevent querying unauthorized tables Detect PII Prevent prompt injection Enforce tenant/user permissions Block unsupported questions Force answers to cite their source Return “I don’t know” when confidence is too low Should guardrails be part of evaluation, inference, or both? Production monitoring Once this is live, what should we log for every AI request? For example: User question → retrieved context → generated SQL/tool calls → query result → final response → model → prompt version → latency → tokens → cost → evaluation scores → user feedback Is that roughly the right model? Regression testing When we change: System prompt Model Retrieval strategy SQL generation logic Tools Temperature Data sources how do you automatically run the benchmark again and determine whether the new version is actually better? Do you set minimum score thresholds before allowing something to deploy? Human feedback How are people incorporating thumbs-up/down or analyst review into their evaluation datasets? Do production failures automatically become new benchmark cases? Making answers more precise This is ultimately my main goal. If our AI currently gives an answer that is “mostly correct,” what is the systematic process for figuring out why it isn’t fully correct? Is the best workflow something like: Production traces → identify failure → categorize failure → add to benchmark → improve retrieval/prompt/tool → run eval → compare against baseline → deploy → monitor Or is there a better approach? I’m especially interested in what the ideal Databricks-native architecture looks like: User Question → Agent / LLM → Retrieval / SQL / Tools → Databricks data → Response → MLflow tracing → Automated evaluators → Benchmark dataset → Regression testing → Production monitoring If you’ve implemented something like this, I’d love to know what you would build first, second, third, etc. Even a practical example like: Week 1: create benchmark Week 2: add tracing + scorers Week 3: add regression tests Week 4: add guardr […truncated]

00New_Championship39291mo ago
Databricks CommunityCommunity Articles

Solution Accelerator Series | Building Common Sense Product Recommendations With LLMs

002mo ago
Databricks CommunityAnnouncements

Solution Accelerator Series | Building a Chatbot With Large Language Models (LLMs)

003mo ago
Databricks CommunityAnnouncements

Solution Accelerator Series | Large Language Models (LLMs) for Customer Service Analytics

003mo ago
RedditHelp

LLMs access to few delta tables inside unity catalog

I want llm to access few tables (not all tables) either through an api endpoint or mcp. Which is the cleanest way? And secure as well Do I create service principles or genie mcp (add only specific tables to genie space)

26aks-7865mo ago
RedditDiscussion

Are LLM agents good at join order optimization?

Everyone’s excited about LLM agents replacing “traditional systems” but what if we pointed them at one of the hardest classical problems in data engineering - SQL join order optimization.? A new blog from Databricks explores exactly that - and the results are both surprising and humbling. Traditionally, query optimizers rely on decades of research, heuristics, and cost models to decide the best join order (because getting it wrong can absolutely destroy performance). So naturally, the question is - **can LLM agents actually do better**? The answer: sometimes yes, but not in the way you might expect. The research shows that LLM agents can: \- Explore join strategies beyond fixed heuristics \- Adapt to specific queries and datasets \- Even outperform built-in optimizers in certain scenarios (\~1.3× improvements reported) LLMs are great at: \- reasoning over complex search spaces \- generating candidate plans \- adapting dynamically While traditional optimizers are still unmatched in: \- consistency \- guarantees \- efficiency at scale So, the future isn’t “LLMs vs systems” - it’s hybrid systems. [Are LLM agents good at join order optimization? | Databricks Blog](https://www.databricks.com/blog/are-llm-agents-good-join-order-optimization)

51szymon_dybczak5mo ago
Stack Overflow

Connecting to Databricks API (hosted LLM in model-serving) via PAT

I'm running code in my local IDE that connects to Databricks's API to pas text into an LLM that is hosted on Databricks. I'm using Personal Access Tokens to get up and running quickly. I'm able to get it working when I add "all scopes" to the PAT, but that is WAAY too much access, and I want to give it just the right access for the task it needs. BUT, I can't figure out which scopes are actually needed. Additional Context: The app is written in Python. It retrieves text from the open internet and then formats it. I want to use an LLM to summarize the content. I would prefer to use an LLM that is hosted on Databricks, via the Databricks API (as opposed to, for instance, using the OpenAI API or Claude API) because I want to use multiple APIs and evaluate them, which Databricks allows. Some of the things I have tried. Databricks lists what the different scopes are and what they do here .: mlflow and model-serving clusters, commang-execution, custom-llms, dashboards, dataclassification, dataquality, environments, files, forecasting, genie, global-init-scripts, instance-pools, instance-profiles, jobs, knowledge-assistants, libraries, mlflow, model-serving, notifications, pipelines apps, clusters, custom-llms, dataclassification, files, genie, global-init-scripts, jobs, knowledge-assistants, marketplace, mlflow, model-serving, secrets, sql, unity-catalog, workspace Initially, I used the OpenAI SDK (as recommended by an LLM), and then switched to the Databricks SDK (because I had hoe that would resolve my issues). Currently, my requirements.txt is: python-dotenv>=1.0.0 openai>=1.0.0 databricks-sdk>=0.49.0

pythondatabricksazure-databrickslarge-language-modelbest-practices
01Mickle-The-Pickle5mo ago
HackerNews

Show HN: An unstructured data workspace for data transformations with LLM

hi HN! a couple of months ago I had to analyze a few thousand audio recordings to help identify issues with customer support. i was able to get some raw high-level initial results with python scripts invoking LLM APIs, but they were too general and unhelpful. writing basic prompts is easy, but tuning them and making them specific enough to ensure no faint signal is missed is hard. you need to iterate through the data with an initial prompt, segment the data into different buckets, chain another prompt for each bucket etc. Then you need to constantly review the raw data to tweak the prompts just the right way to get the desired results. There are no good user-facing tools for scaling to thousands of rows of unstructured data analysis with LLMs. Claude Cowork / agents with access to filesystems are scratching the surface, but having a text-only UI is challenging, especially when you want to go back and adjust your research pipeline, narrow down deterministically to a specific subset of your data with SQL-like filters, or do any cost management. Scaling past 100 files is not well supported. Deep research is difficult to steer and verify. I needed a mini-data warehouse that could help me get insights out of my data, optimize costs with bulk LLM operations (via cost estimation and model choice), and let me browse and verify the data in a user-friendly way, without requiring me to set up something like Databricks. So, I built folio. Folio is a free, local, macOS app for analyzing your unstructured data with LLMs. It's a UI wrapper around a minimal data warehouse that lets users (and agents) do LLM-based transformations on big unstructured datasets. All you need to get started is an AI API key and an account with modal.com Users bring their files into Folio which then get loaded into a tabl, where each row contains a markdown representation of the file contents. Users can then run LLM operations in bulk on those files and use sql filters to create views and narrow down the scope of the transformations. Agents are a first-class citizen and they can plug into folio to do most of the work for you. To take load off the desktop for OCR/Audio Transcription as well as the thousands of http requests to AI APIs we integrate with modal.com as the execution engine. A local orchestrator fans out jobs to modal and then fans them in once complete. Data is never stored anywhere, and only moves in transit through AI API provider and the user's own modal infrastructure. folio workspaces are multi-modal (you can load different data types in the same workspace and move it through the same analysis pipeline) and they can support thousands of files. People use folio today to: - review customer support tickets/emails: bucket issue into different categories, narrow in on categories of interest, and then action that data by generating a response. - extract detailed data from financial documents: load all data that can be found on a particular company, extract structured data like revenue numbers and projections. - do literature reviews: there are lots of agents that help you load data from research paper repositories. once that data is loaded into folio, users can do a steerable deep research over those files. - perform criteria-based search: generate yes/no criteria like "document contains data on XYZ", "document mentions ABC", "documented cites XYZ". Companies like v7labs, hebbia, Legora, Harvey have similar "Tabular Document Review" features, but they are not scalable or compatible with outside agents like Claude Code. Additionally they require expensive enterprise contracts. I see folio moving beyond data analysis into the perfect companion for agentic tasks that require a human-facing UI/UX, cost management and actioning on data in bulk. Website: https://www.usefolio.ai Github: https://github.com/usefolio/folio X: https://x.com/usefolio_ai Looking forward to hearing what people think!

40nibab6mo ago

Get Tuesday's version of this

Tracking LLM? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first