Skip to content
All topics

Fine-tuning

Recent items mentioning Fine-tuning across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

16 recent items2 news12 videos2 community threads
Reddit

Laya off the benchmark: can a zero-shot decision model route real SQL traffic?

Weekend-ish experiment on Databricks. One table, TPC-H orders, 15M rows, living in two places at once: Lakebase (Databricks' managed Postgres), a continuously synced copy with a btree index on the key. Sub-100ms point lookups, useless for a GROUP BY over 15M rows. Delta behind a serverless SQL Warehouse. Great at scans and aggregations, slow at fetching one row. The synced table is the nice part: native continuous Delta to Lakebase Postgres (needs a PK + CDF), so it's one dataset under one Unity Catalog, replication handled by the platform instead of a homemade pipeline. Then I put laya (convaiinnovations/laya, a non-autoregressive zero-shot decision model, vanilla, no fine-tuning) behind a FastAPI Databricks App. It reads each query and picks the engine. MLflow traces input to decision to execution. I submitted by /u/Limp-Park7849 [link] [comments]

00Limp-Park78491w ago
Reddit

Databricks Serverless Compute: Practical Recommendations, Cost-Performance Guidance, and When to Use Serverless Instead of SQL Serverless Warehouses or Classic Clusters

Databricks Serverless compute is a snappy choice for every-day queries with some cost trade-offs. For ETL workloads, it is both fast + cost-efficient. In either scenario and unlike common misconception, there is no cluster fine-tuning needed. Some more detailed takeaways from some testing on my coffee dataset + practical guidance below. Quick disclaimer: I work for a Databricks aligned consulting company. All reporting and recommendations here in any direction are independently made. General Observations from Performance Testing 1) Speed wise, much to my surprise, Serverless beat some of my previous Databricks’ SQL Serverless performance testing. Not across all queries, but in nearly half. 2) Databricks’ SQL Serverless compute is still the best compute option from a pure cost-perspective when it comes to “every day” analytics needs, but from an ETL perspective, “Serverless” is very comparable, without the trial & error of even doing t-shirt sizing, let alone knobs. 3) The scaling time to support an initial large workload was noticeable on the first query and does have room for improvement. I did a dummy print() statement right before (which took like 1-2 secs to warm up). Then hitting that first query containing the joins and 7B+ rows fact table clearly gave the engine something to chew on. For reference, on my SQL Serverless test a while back, the same query took about 20 secs on a large, 15 on XL. Practical Serverless Recommendations 1) If I was orchestrating ETL workloads with SQL and/or Python, I’d use Serverless compute as the cost of Serverless for jobs is pretty darn affordable. My concerns around costs are completely negated when it comes to jobs compute. For new ETL workloads, Serverless is an easy choice for me. For existing ETL workloads, I’d try out to see if Serverless lowers your costs and/or improves performance. Also explore any potential compatibility issues. If you have heavily fine-tuned workloads, that don’t really require much maintenance, migrating the workload might not be best. 2) For daily Python work on established Databricks environments, I would consider trying Serverless. Not having to wait for clusters to boot up and fine-tuning is likely going to lead to productivity gains offsetting any higher costs from interactive Serverless compute itself. No more “left the cluster on by accident” or “I oversized it” type of situations. That said, there are orgs with very experienced folks where cluster management feels like second nature to them. 3) If I was exclusively working with analytical SQL workloads, I’d still use Databricks’ SQL Serverless compute for my day to day queries + dashboards. 4) For folks new to Databricks, I would generally stick to “Serverless” compute all around, and use the “SQL Serverless” warehouses only for BI type experiences. Closing Thoughts on Serverless Having very performant compute without worrying about any configurations (other than rate limits to give you cost controls if you wish) is quite nice from a productivity standpoint. No node count, DBR versioning, t-shirt sizing, nothing. You have plenty of compute options, with or without fine-tuning ! I stress this "no fine-tuning" because there is a plethora of LLM generated articles outdated on Databricks' compute options. Thank you for reading! *No AI was used to write or assist in writing this post. For better or for worse. **Dataset and queries used, configured for the 7B+ row count: https://github.com/joshbogran/coffeeshopdatageneratorv2 submitted by /u/JosueBogran [link] [comments]

00JosueBogran2w ago

Get Tuesday's version of this

Tracking Fine-tuning? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first