Enterprise Cost Optimization & Governance for Gaming Platform
FanDuel, one of America's leading sports gaming and entertainment companies, had seen its Databricks environment grow organically alongside rapid business expansion. Without standardized resource management, monthly Databricks costs had ballooned to $300K with limited visibility into which teams, workloads, or projects were driving spending. Uncontrolled cluster sprawl, inconsistent resource allocation, and inadequate role-based access controls created both financial and governance challenges. Vrahad Analytics was engaged to implement a comprehensive cost optimization and governance strategy that would bring costs under control while maintaining — and even improving — platform performance for all user personas.
Client
FanDuel
Cloud Platform
AWS
Duration
3 Months
The Challenge
FanDuel's Databricks environment reflected the classic pattern of organic growth outpacing governance. As the company scaled from startup to industry leader, the data platform grew without the guardrails needed to manage costs and access at enterprise scale.
Monthly Databricks compute costs had reached $300K and were trending upward, with no clear attribution model to understand which teams, projects, or workloads were driving the majority of spending.
Cluster sprawl was rampant — over 200 active clusters with no standardized sizing, auto-termination policies, or lifecycle management, meaning idle clusters routinely ran 24/7 consuming expensive on-demand compute.
Inconsistent resource allocation meant that data analysts were running notebooks on the same cluster configurations as heavy ETL workloads, leading to both overspending on analyst clusters and performance issues on shared resources.
Limited cost visibility across teams made it impossible for engineering managers to make informed decisions about resource allocation or to hold teams accountable for their Databricks consumption.
Inadequate role-based access controls created security risks — analysts could create arbitrary cluster configurations, and there was no enforcement of organizational standards for compute resources.
No proactive cost monitoring meant that spending spikes (caused by misconfigured jobs, runaway notebooks, or forgotten clusters) were only discovered at the end of the billing cycle, weeks after the waste occurred.
Our Solution
We implemented a comprehensive cost optimization and governance strategy using Terraform-managed cluster policies as the foundation. The approach balanced cost reduction with user experience, ensuring that each persona — analysts, engineers, and data scientists — received appropriately sized resources with automated guardrails that prevent waste without impeding productivity.
Conducted a thorough cost analysis of the existing environment, profiling every cluster's utilization patterns, idle time, and cost attribution. Identified that 45% of total spend was on idle or significantly underutilized clusters.
Designed and deployed Terraform-managed cluster policies for three distinct user personas: analysts (small, short-lived interactive clusters), data engineers (medium, auto-scaling batch clusters), and data scientists (GPU-enabled, spot-heavy ML clusters).
Enforced auto-termination after 1 hour of inactivity across all interactive clusters via cluster policies, eliminating the single largest source of waste — clusters left running overnight and over weekends by users who forgot to shut them down.
Deployed spot instances for all fault-tolerant workloads (batch ETL, model training, data quality checks), achieving up to 90% compute cost savings on these workloads by leveraging AWS's excess capacity at steep discounts.
Implemented Unity Catalog for unified governance with least-privilege access controls, ensuring that users could only create clusters that conform to their persona's policy and access only the data assets relevant to their role.
Built a continuous cost monitoring and alerting system using custom dashboards that provide daily cost breakdowns by team, workload type, and cluster — with automated Slack alerts for spending anomalies that exceed thresholds.
Implementation Phases
Cost Discovery & Profiling
2 WeeksAnalyzed 90 days of Databricks usage data to profile every cluster's utilization, idle time, and cost. Categorized workloads by type (interactive, batch, ML) and mapped spending to teams and projects. Identified $180K/month in optimization opportunities.
Cluster Policy Design
2 WeeksDesigned three persona-based cluster policies (analyst, engineer, data scientist) with Terraform, defining allowed instance types, auto-scaling ranges, auto-termination settings, spot instance ratios, and tagging requirements for cost attribution.
Terraform Deployment & Policy Enforcement
3 WeeksDeployed cluster policies via Terraform across all workspaces, migrated existing clusters to policy-compliant configurations, and enabled policy enforcement that prevents creation of non-compliant clusters.
Spot Instance Strategy & Optimization
2 WeeksImplemented spot instance pools for fault-tolerant workloads with automatic fallback to on-demand for critical jobs. Configured instance fleet strategies to maximize spot availability while minimizing interruption risk.
Monitoring, Alerting & Governance
3 WeeksBuilt real-time cost monitoring dashboards, implemented automated Slack alerts for spending anomalies, deployed Unity Catalog for data governance, and conducted training sessions for all user personas on the new policies and self-service capabilities.
Technologies Used
Key Results
Measurable outcomes and business impact delivered through this engagement.
Reduced monthly Databricks costs from $300K to $120K — a 60% savings ($180K/month, $2.16M annualized) — while maintaining or improving performance for all workload types.
Implemented role-based cluster policies for 3 distinct user personas (analysts, engineers, data scientists), ensuring each group gets appropriately sized resources without over-provisioning.
Achieved up to 90% compute cost savings on fault-tolerant workloads through strategic spot instance deployment with automatic fallback to on-demand for business-critical jobs.
Established continuous cost monitoring with real-time dashboards and automated anomaly alerting, reducing mean time to detect spending spikes from weeks (end of billing cycle) to minutes.
Enforced least-privilege access and unified governance via Unity Catalog, ensuring that all users operate within organizational guardrails while maintaining self-service capabilities.
Eliminated cluster sprawl entirely — active clusters reduced from 200+ unmanaged to 80 policy-compliant clusters with automated lifecycle management and utilization tracking.
Before vs. After Comparison
Monthly Cost
Before
$300,000
After
$120,000
60% reduction
Annual Savings
Before
$3.6M/year
After
$1.44M/year
$2.16M saved
Active Clusters
Before
200+ unmanaged
After
80 policy-managed
60% fewer
Idle Waste
Before
45% of spend
After
< 5% of spend
90% eliminated
Spot Instance Usage
Before
0%
After
70% of batch
Up to 90% savings
Cost Anomaly Detection
Before
End of month
After
Real-time
Minutes vs weeks
Monthly Databricks Cost ($K)
Compute Cost Distribution (After)
Platform Maturity — Before vs After
Monthly Cost Trend ($K)
Previous Case Study
GMG — SAP HANA to Databricks Migration
Next Case Study
DriveWealth — Legacy to E2 Migration
Facing a Similar Challenge?
Let us help you architect a solution that delivers measurable business impact.
