Databricks
Databricks is a unified analytics platform that brings data engineering, data science and business analytics into a single workspace. Built by the original creators of Apache Spark, it covers large-scale data processing, machine learning workflows and real-time analytics.
We have extensive Databricks knowledge in house, which is why it moves to adopt. It is also an established choice among our clients, so a good part of this work is joining teams where Databricks is already running and helping them get more out of it.
Considerations
Vendor lock-in: Databricks is a commercial platform. Pipelines written against its proprietary features are not trivially portable, and both cost and roadmap sit with a single vendor.
US cloud provider dependency: Databricks runs on the US hyperscalers. For clients with digital sovereignty requirements that dependency has to be weighed explicitly — see European Sovereign Cloud.
Alternatives
Databricks is not the only thing we offer. We also have extensive knowledge of open source lakehouse setups, and for our own projects that is what we prefer: it avoids the lock-in and the sovereignty questions, at the cost of assembling and operating more of the stack ourselves.
For client work we evaluate which of the two is the better fit — what the client already runs, the engineering capacity available to operate it, and the budget all weigh in. Neither is the default answer.
Databricks is a unified analytics platform that accelerates innovation by unifying data science, engineering, and business analytics. Built by the original creators of Apache Spark™, it offers a cloud-based environment for processing large-scale data, running machine learning models, and enabling real-time analytics.
Why Databricks?
Unified Platform: Combines data engineering, data science, and business analytics in one collaborative workspace.
Scalable Processing: Optimized for big data with Apache Spark™, allowing for efficient processing of large datasets.
Collaborative Notebooks: Features interactive notebooks with real-time co-authoring, version control, and support for multiple programming languages.
Machine Learning Integration: Provides seamless integration with popular ML frameworks like TensorFlow, PyTorch, and scikit-learn.
Considerations at INFO
Databricks is being trialed for projects that demand extensive data processing capabilities, advanced analytics, and robust machine learning workflows.
Current Focus: Assessing scalability, ease of collaboration, integration with existing data pipelines, and overall impact on data-driven decision-making.
Databricks' comprehensive analytics and collaborative features make it a promising tool as we trial its potential to enhance our big data and analytics initiatives.