LLM
Recent items mentioning LLM across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.
Databricks launched ai_decide, a new SQL and REST AI Function designed to deliver faster, lower-cost structured decisions on governed data than traditional LLMs 1. Across evaluation and agent workflows, practitioners are adopting tiered LLM judge patterns ahead of Jev's integration into MLflow 3.17 3 and deploying Unity Gateway tracing to eliminate token waste caused by ambiguous LLM inputs and broken agent tool calls 5.
Generated daily from the 5 most recent items mentioning LLM. Click any [N] to jump to the source.
Introducing ai_decide: make fast decisions on your governed data
Databricks has introduced ai_decide, a new AI Function that processes unstructured text to return structured decisions faster and at lower cost than traditional LLMs. Optimized for decision-making rather than text generation, it runs directly on governed data at batch scale in SQL or in real time over REST for tasks like prompt routing, metadata tagging, and agent quality evaluation.
Databricks Micro Apps, App Spaces and Genie App Generator
Databricks launched App Spaces and Serverless Micro Apps at the end of last week (in Beta). Over the weekend, I tested migrating two of my production Databricks Apps to App Spaces. App Migration 1: Blocked by Zero Egress My Data Portfolio Project Creator required internet access to pull data stacks from live job postings. Then, the LLM needs internet access to research for open data sources to use. In standard Databricks Apps, this runs cleanly. In App Spaces, there is zero external internet egress including for LLMs. The app can reach internal workspace resources, but it cannot touch the outside web. If your app relies on third-party APIs, external databases, or web scraping, App Spaces is a non-starter until Databricks opens network egress. Workload 2: Success on Internal FinOps My client-facing DBU cost observability app reads workspace usage data and writes directly to Lakebase. Because it requires zero external network calls, the migration worked. Both the micro app compute and Lakebase scale to zero when idle. Cold starts take roughly 30 seconds (though in beta, you occasionally need a quick browser refresh once it spins up). For internal, low-frequency administrative tools, this turns a continuous monthly compute bill into pennies. (+ App Spaces, Genie App Generator, and Micro Apps are free while in beta) The next part is less about App Spaces and more just general best practice for Databricks Apps that I see people miss. Stop Using Delta Lake as an OLTP Database Databricks Apps are software applications, not batch analytics notebooks. If your app writes application state, session data, or row-level CRUD directly into analytical Delta tables, you need to rethink that design. For app transactions, use Lakebase (serverless Postgres which also scales to 0). Both are governed under UC, but Lakebase gives your app the low-latency transactional engine that application engineering actually requires. My Verdict on the Beta App Spaces solves the idle compute problem that has plagued Databricks Apps since launch. But until Databricks allows us to deploy directly from existing Git repos and opens external network egress, it remains limited to internal-only use cases. Curious on other peoples experience... How has it been for others? submitted by /u/OkImprovement7010 [link] [comments]
Can Jev replace your LLM judge? Part 2: Testing harder answers
TypeSafe's Jev model evaluates MLflow QA answers roughly seven times faster and significantly cheaper than larger LLMs, but it misses subtly incorrect technical details such as reversed roles and omitted permissions. Ahead of Jev integration arriving in MLflow 3.17, practitioners can use its confidence probabilities to fast-track clear verdicts while escalating borderline cases to a larger model for review.
NewsWhat are Agents & How they Work? #agent #llm #genai
AI agents combine a large language model for reasoning, external digital tools for executing actions, and memory for tracking context to complete complex tasks autonomously. This architecture allows the system to continuously loop through planning, executing, and adjusting steps until a defined goal, such as organizing a chaotic folder of files, is fully achieved.
The deployment decalogue
I am an ML engineer, but I come from a software engineering background: years of full-stack work, with heavy DevOps and Terraform experience. I come from teams that deploy to production five times a day with real continuous deployment. And honestly? Pressing the button still feels weird sometimes. Every engineer knows that feeling, no matter how good the safety net is. So I wrote down the list that settles it. Ten commandments, one flow, written with data scientists and ML teams in mind, but it works for batch jobs, realtime inference, and LLMs alike. Answer honestly, and if all ten are true, you can ship to production anytime, in any form or way. submitted by /u/SuspiciousPavement [link] [comments]
How we eliminated $1 million a year of wasted AI agent spend in one hour
Broken MCP tool calls and silent retries can quietly waste over $1 million annually in tokens and engineering hours across AI agent fleets. Tracing tool calls with Unity Gateway and analyzing spend with Genie One enables teams to rapidly deploy fixes and design tools that gracefully handle ambiguous LLM inputs.
How are you actually setting up AI/LLM evals in Databricks end-to-end? Looking for a step-by-step production workflow
We have a product where Databricks is our backend , and we’re now trying to properly evaluate and improve the quality of the AI-generated answers in our application. I’m looking for advice from people who have actually implemented LLM/GenAI evaluations in Databricks in production . I’d really appreciate an end-to-end, step-by-step explanation of how you would set this up from scratch. Specifically, how would you approach: Define what a “good answer” means Accuracy / correctness Relevance Completeness Groundedness / faithfulness to our data Hallucination rate Citation/source correctness Following user instructions Response consistency Latency and cost Create a benchmark / golden dataset Should we manually create a set of representative user questions? How many questions are enough to start? Should each question have an expected answer? Should we store expected SQL/results, expected sources, or just an expected natural-language response? Where should this benchmark dataset live in Databricks? How do you keep it updated as the product evolves? Set up automated evaluations What Databricks/MLflow tools should we be using today? MLflow evaluation? LLM-as-a-judge? Custom scorers? Human evaluation? How do you combine these rather than relying on one score? Evaluate RAG / data-grounded answers Our AI answers questions based on enterprise data in Databricks. How do you separately evaluate: Retrieval quality Whether the correct tables/documents were selected Context relevance Groundedness Final answer correctness Whether the model invented something that wasn’t in the retrieved data Evaluate text-to-SQL / analytics questions If a user asks something like: “What was average building occupancy last month?” should we evaluate: Generated SQL Tables selected SQL execution result Final natural-language answer separately? What is the recommended architecture for this? Guardrails Where should guardrails sit in the architecture? For example: Prevent hallucinated numbers Prevent querying unauthorized tables Detect PII Prevent prompt injection Enforce tenant/user permissions Block unsupported questions Force answers to cite their source Return “I don’t know” when confidence is too low Should guardrails be part of evaluation, inference, or both? Production monitoring Once this is live, what should we log for every AI request? For example: User question → retrieved context → generated SQL/tool calls → query result → final response → model → prompt version → latency → tokens → cost → evaluation scores → user feedback Is that roughly the right model? Regression testing When we change: System prompt Model Retrieval strategy SQL generation logic Tools Temperature Data sources how do you automatically run the benchmark again and determine whether the new version is actually better? Do you set minimum score thresholds before allowing something to deploy? Human feedback How are people incorporating thumbs-up/down or analyst review into their evaluation datasets? Do production failures automatically become new benchmark cases? Making answers more precise This is ultimately my main goal. If our AI currently gives an answer that is “mostly correct,” what is the systematic process for figuring out why it isn’t fully correct? Is the best workflow something like: Production traces → identify failure → categorize failure → add to benchmark → improve retrieval/prompt/tool → run eval → compare against baseline → deploy → monitor Or is there a better approach? I’m especially interested in what the ideal Databricks-native architecture looks like: User Question → Agent / LLM → Retrieval / SQL / Tools → Databricks data → Response → MLflow tracing → Automated evaluators → Benchmark dataset → Regression testing → Production monitoring If you’ve implemented something like this, I’d love to know what you would build first, second, third, etc. Even a practical example like: Week 1: create benchmark Week 2: add tracing + scorers Week 3: add regression tests Week 4: add guardr […truncated]
EventsFrom Primitives to Production: How Anthropic Builds Agents
Anthropic defines agents as LLMs in loops with tool access, leaning on model intelligence over rigid workflows, and uses domain-specific skills, MCPs, and evals to build effective agents across sectors. A demo showed a site reliability agent autonomously identifying a database pool issue, fixing the code, and generating a postmortem.
NewsAI Runtime CLI | Serverless GPU LLM Training
Databricks AI Runtime is a CLI tool that enables distributed LLM training on serverless GPUs using YAML configuration files, supporting deployments up to 256 H100 GPUs with built-in MLflow experiment tracking and hardware metrics. Users develop training code locally in their preferred environment and submit jobs via CLI commands that automatically handle infrastructure provisioning, log streaming, and performance monitoring.
TutorialsDetect Energy Theft Faster with Genie
Databricks demonstrates an end-to-end AI application that detects energy theft, automates investigations, and generates executive reports using Unity Catalog and Genie. The video walks through an architecture featuring Lakebase for transactional storage, model serving for machine learning and LLMs, and AI gateways for governance and cost control.
NewsRAG Explained + Build a RAG App From Scratch in Python using LLM | Chapter 08
This video teaches the core concepts of retrieval-augmented generation and demonstrates how to build a complete RAG application from scratch in Python using a Groq language model. The tutorial covers a seven-step pipeline including document ingestion, token-based text chunking, vector embedding generation, in-memory storage, similarity search, prompt augmentation, and response generation.
AI-Enabled Advisory Services for Higher Education
Databricks has introduced a GenAI-powered solution that automates and scales the quality review of higher education advisory calls by transcribing conversations, scoring performance against institutional rubrics, and surfacing insights. The blog post demonstrates this architecture on a single, governed platform, providing code notebooks and a Genie space for natural language data exploration.
Solution Accelerator Series | Building Common Sense Product Recommendations With LLMs
NewsAI-Ready Data on Databricks: How TK Elevator Uses Context and Meaning to Make AI Agents Work
TK Elevator uses Databricks' Unity Catalog to create "AI-ready data" by harmonizing data from over 100 disparate systems, enabling a common language for their AI agents. This foundation supports predictive maintenance for elevators, empowering 25,000 service technicians with tailored support and voice debriefing capabilities.
Guide to Agentic Systems and AI Agents
Agentic AI systems are autonomous software platforms that perceive, reason, execute multi-step tasks, and learn with minimal human intervention, unlike traditional generative models. These systems use LLMs as reasoning engines with external tools and memory to complete complex workflows, with enterprise adoption spanning customer service to financial risk.
End-to-End RAG Workflow: How Retrieval Augmented Generation Works
Databricks now offers a five-stage RAG workflow for connecting LLMs to external knowledge bases, enabling accurate, domain-specific answers without model retraining. Production RAG requires careful selection of embedding models, vector database indexing, chunking strategies, and hybrid search, with independent evaluation of retrieval precision and generation faithfulness.
NewsAn agentic CDP empowers marketing teams to move from static campaigns to Infinity Campaigns
Databricks introduces "Infinity Campaigns," a new marketing approach where an AI agent processes customer signals to determine the next best action and generate personalized content. This creates a continuous feedback loop where customer interactions generate new signals, feeding back into the agent for ongoing optimization.
EventsIntroducing CustomerLake: The Agentic CDP | Ali Ghodsi Databricks CEO at Data + AI Summit
Databricks introduces CustomerLake, an agentic Customer Data Platform built on the lakehouse architecture. It features a profile agent for identity deduplication using LLMs and a campaign agent for personalized, one-to-one "infinity campaigns."
Solution Accelerator Series | Building a Chatbot With Large Language Models (LLMs)
What is document AI?
Document AI transforms messy, high-volume documents into structured data for downstream systems, offering value beyond just faster processing. While generative AI makes it more adaptable for summarization and extraction, accuracy still relies on validation and human review, with governance becoming central due to sensitive data.
How Ecolab rebuilt retail intelligence on Databricks and Anthropic Claude
Ecolab rebuilt retail intelligence on Databricks and Anthropic Claude, converting 700-page FDA manuals into real-time answers for frontline staff using Foundation Model APIs and cutting compliance report compilation from two weeks to under two minutes. The solution, a native Databricks App with Lakebase Postgres and Unity Catalog, unifies nine siloed data sources and employs a multi-agent orchestration framework with Judge LLMs and MLflow tracing for personalized, continuously refined intelligence.
AI Serving Platform That Adapts to Your Model
Databricks now offers a fully managed AI serving platform that automatically adapts to your model's resource needs, from scikit-learn to 70B LLMs, without manual configuration. This results in up to 90% lower infrastructure costs and <10ms p99 latency overhead for customers migrating from self-managed stacks.
Announcing the Databricks storage ecosystem: Governing the enterprise data estate, wherever it lives
The Databricks Storage Ecosystem now natively connects hybrid and on-premises storage platforms to Databricks via OpenSharing, enabling centralized data governance and GenAI scaling across your entire hybrid infrastructure. Run Databricks Serverless Compute, Genie, and LLMs directly on your on-premises datasets with a zero-copy architecture, instantly turning isolated data into active, AI-ready assets.
Solution Accelerator Series | Large Language Models (LLMs) for Customer Service Analytics
NewsHow LLMs Understand your Prompts: Tokenization & Embeddings | Chapter 05
The video explains how Large Language Models (LLMs) understand text by converting it into numerical representations through tokenization and embeddings. It demonstrates how text is broken into tokens, assigned unique IDs, and then transformed into dense vectors (embeddings) that capture semantic meaning and positional information for LLM processing.
NewsWhen to choose CPU vs GPU: Databricks AI Runtime Explained
CPUs are best for data work like ETL, feature engineering, SQL, and classical machine learning, while GPUs are designed for deep learning workloads such as fine-tuning LLMs and training neural networks. Databricks AI Runtime simplifies GPU usage by providing serverless Nvidia GPUs, removing the need for manual infrastructure setup and allowing seamless transitions between CPU for data prep and GPU for model training within the Databricks environment.
TutorialsHow Large Language Models (LLMs) Work - Full Explanation | Chapter 04
Large Language Models (LLMs) are text-based neural networks trained on massive data to predict the next word (token), operating through tokenization, vector embeddings, and a transformer architecture. LLMs undergo pre-training, supervised fine-tuning, and reinforcement learning from human feedback to become helpful, safe, and aligned, with concepts like context length, knowledge cut-off, and hallucination defining their capabilities and limitations.
Accelerating LLM Inference with Prompt Caching for Open‑Source Models on Databricks
Databricks now supports prompt caching for open-source models across all workloads, automatically accelerating LLM inference by reusing repeated prompt prefixes. This feature boosts throughput by 2.5x and reduces P50 latency by 3x for models like GPT-OSS, with no setup required.
TutorialsBuilding Trustworthy, High-Quality AI Agents with MLflow
Databricks' MLflow platform helps developers build trustworthy, high-quality AI agents by providing tools for end-to-end observability, evaluation, prompt management, and AI gateway governance. It demonstrates how MLflow facilitates tracing, expert feedback collection, automated issue detection with LLM judges, prompt optimization, and continuous monitoring throughout the agent development lifecycle.
TutorialsAI Agents That Remember: Building Stateful Systems with Lakebase
AI agents require four types of memory (working, episodic, entity, procedural) to be truly intelligent and stateful, which traditional databases struggle to provide. Databricks Lakebase, built on Postgres, offers a unified OLTP and OLAP solution with features like serverless auto-scaling and Git-style branching to manage these complex memory needs for AI agents.
NewsGovern MCP servers in Databricks #databricks #mcp #aigovernance
Databricks Unity AI Gateway now governs MCP servers, centralizing their management alongside built-in foundation models and LLMs. This integration allows for easier governance and orchestration of various AI components and agents within Databricks.
TutorialsHow to Build an AI Security Governance Hub with Agent Bricks
Databricks Agent Bricks enables building an AI Security Governance Hub by transforming static security playbooks into adaptive multi-agent systems. The video demonstrates combining a knowledge assistant for unstructured documents and a Genie space for structured data into a supervisor agent, then details how to tune and monitor these agents for improved performance and data privacy.
EventsBuilding Trustworthy, High-Quality AI Agents with MLflow
MLflow provides a comprehensive platform for building, evaluating, and deploying high-quality AI agents, offering tools for observability, automated evaluation, prompt optimization, and production monitoring. It enables developers to streamline the agent development lifecycle, from prototyping and testing with human and AI judges to fixing issues and ensuring reliable, governed deployment.
LLMs access to few delta tables inside unity catalog
I want llm to access few tables (not all tables) either through an api endpoint or mcp. Which is the cleanest way? And secure as well Do I create service principles or genie mcp (add only specific tables to genie space)
LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools
LLMs are a subset of AI, and this guide clarifies their practical differences, use cases, and tools. Understand how LLMs fit into the broader AI landscape and what that means for your Databricks workflows.
AI Applications: Tools, Use Cases, and Platforms
AI applications span four capability tiers, each with distinct data requirements and evaluation frameworks, and enterprise deployments often stall due to inadequate data infrastructure. Production-grade model development, from prompt engineering to pretraining, is increasingly accessible with open-source LLMs, but requires pre-built governance and monitoring infrastructure for successful deployment at scale.
NewsFrom AI to Agents| Fundamentals of AI | ML | DL | LLM & GenAI | Chapter 01
The video explains the fundamental concepts of AI, ML, DL, LLMs, and GenAI, illustrating their hierarchical relationship as subsets of each other. It also defines what models are (mathematical formulas trained on data) and how agents combine LLMs with tools and optional memory to perform autonomous tasks.
Are LLM agents good at join order optimization?
Everyone’s excited about LLM agents replacing “traditional systems” but what if we pointed them at one of the hardest classical problems in data engineering - SQL join order optimization.? A new blog from Databricks explores exactly that - and the results are both surprising and humbling. Traditionally, query optimizers rely on decades of research, heuristics, and cost models to decide the best join order (because getting it wrong can absolutely destroy performance). So naturally, the question is - **can LLM agents actually do better**? The answer: sometimes yes, but not in the way you might expect. The research shows that LLM agents can: \- Explore join strategies beyond fixed heuristics \- Adapt to specific queries and datasets \- Even outperform built-in optimizers in certain scenarios (\~1.3× improvements reported) LLMs are great at: \- reasoning over complex search spaces \- generating candidate plans \- adapting dynamically While traditional optimizers are still unmatched in: \- consistency \- guarantees \- efficiency at scale So, the future isn’t “LLMs vs systems” - it’s hybrid systems. [Are LLM agents good at join order optimization? | Databricks Blog](https://www.databricks.com/blog/are-llm-agents-good-join-order-optimization)
Are LLM agents good at join order optimization?
LLM agents can improve Databricks join order optimization, achieving 1.3x latency reduction in 80% of cases by reasoning through runtime statistics. This prototype demonstrates LLM agents' potential to act as data-driven DBAs, addressing cardinality misestimation challenges in complex SQL queries.
NewsGenAI - For Data Engineers Agenda & Introduction | LLM & Agentic AI | LangChain & LangGraph | Claude
This video introduces a new course, "GenAI for Data Engineers," designed to teach data engineers how to leverage generative AI, LLMs, and agentic AI. The course covers basics of LLMs, building agents with LangChain and LangGraph, using Cloud Code, and applying agentic AI within Databricks and data engineering workflows.
Connecting to Databricks API (hosted LLM in model-serving) via PAT
I'm running code in my local IDE that connects to Databricks's API to pas text into an LLM that is hosted on Databricks. I'm using Personal Access Tokens to get up and running quickly. I'm able to get it working when I add "all scopes" to the PAT, but that is WAAY too much access, and I want to give it just the right access for the task it needs. BUT, I can't figure out which scopes are actually needed. Additional Context: The app is written in Python. It retrieves text from the open internet and then formats it. I want to use an LLM to summarize the content. I would prefer to use an LLM that is hosted on Databricks, via the Databricks API (as opposed to, for instance, using the OpenAI API or Claude API) because I want to use multiple APIs and evaluate them, which Databricks allows. Some of the things I have tried. Databricks lists what the different scopes are and what they do here .: mlflow and model-serving clusters, commang-execution, custom-llms, dashboards, dataclassification, dataquality, environments, files, forecasting, genie, global-init-scripts, instance-pools, instance-profiles, jobs, knowledge-assistants, libraries, mlflow, model-serving, notifications, pipelines apps, clusters, custom-llms, dataclassification, files, genie, global-init-scripts, jobs, knowledge-assistants, marketplace, mlflow, model-serving, secrets, sql, unity-catalog, workspace Initially, I used the OpenAI SDK (as recommended by an LLM), and then switched to the Databricks SDK (because I had hoe that would resolve my issues). Currently, my requirements.txt is: python-dotenv>=1.0.0 openai>=1.0.0 databricks-sdk>=0.49.0
Show HN: An unstructured data workspace for data transformations with LLM
hi HN! a couple of months ago I had to analyze a few thousand audio recordings to help identify issues with customer support. i was able to get some raw high-level initial results with python scripts invoking LLM APIs, but they were too general and unhelpful. writing basic prompts is easy, but tuning them and making them specific enough to ensure no faint signal is missed is hard. you need to iterate through the data with an initial prompt, segment the data into different buckets, chain another prompt for each bucket etc. Then you need to constantly review the raw data to tweak the prompts just the right way to get the desired results. There are no good user-facing tools for scaling to thousands of rows of unstructured data analysis with LLMs. Claude Cowork / agents with access to filesystems are scratching the surface, but having a text-only UI is challenging, especially when you want to go back and adjust your research pipeline, narrow down deterministically to a specific subset of your data with SQL-like filters, or do any cost management. Scaling past 100 files is not well supported. Deep research is difficult to steer and verify. I needed a mini-data warehouse that could help me get insights out of my data, optimize costs with bulk LLM operations (via cost estimation and model choice), and let me browse and verify the data in a user-friendly way, without requiring me to set up something like Databricks. So, I built folio. Folio is a free, local, macOS app for analyzing your unstructured data with LLMs. It's a UI wrapper around a minimal data warehouse that lets users (and agents) do LLM-based transformations on big unstructured datasets. All you need to get started is an AI API key and an account with modal.com Users bring their files into Folio which then get loaded into a tabl, where each row contains a markdown representation of the file contents. Users can then run LLM operations in bulk on those files and use sql filters to create views and narrow down the scope of the transformations. Agents are a first-class citizen and they can plug into folio to do most of the work for you. To take load off the desktop for OCR/Audio Transcription as well as the thousands of http requests to AI APIs we integrate with modal.com as the execution engine. A local orchestrator fans out jobs to modal and then fans them in once complete. Data is never stored anywhere, and only moves in transit through AI API provider and the user's own modal infrastructure. folio workspaces are multi-modal (you can load different data types in the same workspace and move it through the same analysis pipeline) and they can support thousands of files. People use folio today to: - review customer support tickets/emails: bucket issue into different categories, narrow in on categories of interest, and then action that data by generating a response. - extract detailed data from financial documents: load all data that can be found on a particular company, extract structured data like revenue numbers and projections. - do literature reviews: there are lots of agents that help you load data from research paper repositories. once that data is loaded into folio, users can do a steerable deep research over those files. - perform criteria-based search: generate yes/no criteria like "document contains data on XYZ", "document mentions ABC", "documented cites XYZ". Companies like v7labs, hebbia, Legora, Harvey have similar "Tabular Document Review" features, but they are not scalable or compatible with outside agents like Claude Code. Additionally they require expensive enterprise contracts. I see folio moving beyond data analysis into the perfect companion for agentic tasks that require a human-facing UI/UX, cost management and actioning on data in bulk. Website: https://www.usefolio.ai Github: https://github.com/usefolio/folio X: https://x.com/usefolio_ai Looking forward to hearing what people think!
Introducing MLflow AI Gateway: Governed, Observable Access to LLMs
MLflow AI Gateway provides a single, secure endpoint for all LLM providers, complete with usage tracking and native tracing. This new feature offers governed, observable access to LLMs for Databricks practitioners.
MemAlign: Building Better LLM Judges From Human Feedback With Scalable Memory
MemAlign, a new framework for aligning LLMs with human feedback, is now available, offering competitive or better quality than state-of-the-art prompt optimizers at significantly lower cost and latency. It achieves this through a lightweight dual-memory system, making it a valuable tool for building better LLM judges.
Introducing DeepEval, RAGAS, and Phoenix Judges in MLflow
MLflow has integrated over 50 LLM-as-a-judge metrics from DeepEval, RAGAS, and Arize Phoenix directly into the MLflow Scorer API. Practitioners can now run and compare these third-party evaluations alongside native MLflow judges in a single API call and unified visual UI to debug and evaluate RAG pipelines and agents.
NewsVibe-Engineering LakeFlow Pipelines, the Advancing Analytics Way
Advancing Analytics introduces Lake Forge, an engineering framework that uses LLMs and an agentic workflow to generate standardized LakeFlow pipeline templates from data specifications. This system aims to enable scalable, repeatable, and supportable data pipeline creation by balancing AI-driven "vibe coding" with human-engineered guardrails and validation loops.
NewsViews on upcoming LLM and GenAI Playlist #llm #shorts
The creator is launching a new playlist on AI, LLMs, and Gen AI for data engineers, starting with neural network basics. They seek viewer feedback on whether to offer it free or as a paid subscription due to piracy concerns.
NewsThe Hitchhiker's Guide to Delta Lake Streaming in an Agentic Universe
Configuration-driven development abstracts Delta Lake streaming boilerplate so engineers focus on transformations. Structured API tools with descriptive metadata enable LLMs to automatically generate working Spark streaming pipelines by exploring Unity Catalog through conversation.
NewsAI Agents in Action: Structuring Unstructured Data on Demand With Databricks and Unstructured
Unstructured's MCP server enables LLMs to directly transform unstructured data (PDFs, documents, images) into JSON using natural language commands instead of code. Their agentic approach retrieves data on-demand from sources like S3 and SharePoint, reducing costs compared to traditional full data ingestion into vector databases.
NewsAutomating Engineering with AI - LLMs in Metadata Driven Frameworks
Data engineers must adapt to the AI revolution by using artificial intelligence to automate time consuming tasks like pipeline generation and data cleansing. Professionals should integrate AI coding assistants and metadata driven frameworks into their workflows to increase efficiency and remain competitive.
TutorialsSponsored by: West Monroe | Disruptive Forces: LLMs and the New Age of Data Engineering
LLMs provide three key superpowers for data engineering: universal code migration, automated data model generation, and intelligent documentation, reducing projects from months to hours. Databricks' Hopper app demonstrates this by auto-generating complete data models, ETL mappings, and SQL scripts in minutes from business requirements and source schemas.
EventsJamie Dimon Chairman and CEO of JPMorgan Chase at Data + AI Summit
JPMorgan Chase invests $2 billion annually on AI and has elevated AI/data leadership to the executive management level to drive organizational adoption, enabling 600+ active use cases across all business functions. The primary challenge is consolidating fragmented data across the organization and achieving buy-in from business leaders to implement AI solutions, rather than building the models themselves.
Unity Catalog AI 0.2.0
Unity Catalog AI 0.2.0 adds Gemini and LiteLLM integrations for using UC functions as AI tools and introduces function wrapping APIs with dependency management support. AutoGen integration now requires version 0.4.x and Databricks client now only supports serverless endpoints.
NewsSponsored: EY | Business Value Unleashed: Real-World Accelerating AI & Data-Centric Transformation
Ernst and Young demonstrates two real world databricks use cases for accelerating AI and data transformation, including a retail inventory optimization and demand forecasting model that saved a client over one hundred million dollars in working capital. The presentation also highlights an emerging generative AI solution that processes satellite imagery to make geospatial data searchable through natural language for business users.
NewsAI Regulation is Coming: The EU AI Act and How Databricks Can Help with Compliance
The video outlines the risk-based approach of the upcoming EU AI Act, detailing its compliance obligations, self-conformity assessments, and challenges regarding foundation models. It also highlights Databricks platform tools, such as Unity catalog, ML flow, and lake house AI, designed to assist with AI governance and regulatory compliance.
NewsIFC's MALENA Provides Analytics for ESG Reviews in Emerging Markets Using NLP and LLMs
The International Finance Corporation developed Malena, an artificial intelligence tool using natural language processing and large language models on the Databricks lakehouse platform to automate environmental, social, and governance reviews. The system analyzes lengthy corporate documents and news reports to extract risk terms, assess sentiment, and generate sustainability performance dashboards for emerging market investments.
Get Tuesday's version of this
Tracking LLM? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.




