
Work Here?
GetDBT.com provides a cloud-based data management platform that helps businesses accelerate their data development lifecycle. It lets teams write business logic faster, improve code reuse, and onboard new data developers using a declarative code style, modular components, macros, and thorough documentation of data products. The platform also enforces governance and data quality by enabling testing, proactive issue fixes, and visibility into analytics code and compute spend. Built for scale, it supports team unification, standardized processes, extensible workflows, and a broad set of integrations and APIs from a strong partner ecosystem. The service is offered on a subscription basis with a 14-day free trial, uptime guarantees, and 24x7 support to ensure reliable delivery. The goal is to help data teams mature, reduce development time, improve data quality and governance, and manage costs as data platforms scale.
Industries
Data & Analytics
Enterprise Software
Company Size
1,001-5,000
Company Stage
Series D
Total Funding
$414.4M
Headquarters
Philadelphia, Pennsylvania
Founded
2016
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$414.4M
Above
Industry Average
Funded Over
4 Rounds
Industry standards
Unlimited Paid Time Off
401(k) Company Match
401(k) Retirement Plan
Health Insurance
Paid Parental Leave
Wellness Program
Home Office Stipend
Fivetran and dbt Labs have launched new data infrastructure capabilities aimed at making enterprise data ready for AI agents. The company announced the general availability of dbt v2, a Rust rewrite that parses large projects up to 10 times faster than its predecessor, and dbt State, which optimises warehouse compute by tracking changes through metadata. The firms also introduced Fivetran Context Layer in private beta, which unifies structured and unstructured data to provide context for AI agents. Lake Compute, now in private beta, enables users to run dbt models against Apache Iceberg tables. RxBenefits reported reducing warehouse costs by 59% after deploying dbt State, saving over $8,000 in 60 days. Virgin Media O2 achieved 25% savings on job run time and compute costs.
How Hex thinks shared context will unlock enterprise AI ROI. As enterprises move beyond AI pilots, shared context and open infrastructure will determine whether intelligent agents deliver measurable business value at scale. Smruthi Nadig · AUGUST 21, 2026, 10:42 AM · Updated AUGUST 21, 2026, 12:32 PM Key Takeaways What actually matters. Organisations are shifting focus from AI model selection to demonstrating measurable business outcomes. A shared context and open infrastructure are crucial for scaling enterprise AI effectively. The enterprise AI conversation has changed. Over the past two years, organisations have invested heavily in large language models, copilots and generative AI applications, driven by the promise of greater productivity and faster decision-making. Today, however, boardrooms are asking a different question. Rather than debating which model to adopt, leaders want evidence that AI is delivering measurable business outcomes. That shift took centre stage at MachineCon USA on July 24, where Carlos Aguilar, Head of Product at Hex, joined leaders from Fivetran, dbt Labs and other enterprise data teams for a panel titled 'The AI-Native Mandate: Building the Foundation for Measurable AI ROI'. The discussion reflects a growing consensus across the industry: organisations can no longer think of AI as a standalone capability. They need a foundation that enables people and AI systems to reason, collaborate and make decisions from the same trusted understanding of the business. Hex believes the biggest obstacle to scaling enterprise AI is no longer the model itself, but the context surrounding it. As Aguilar puts it, "Models are commoditising fast. Context isn't. Every agent, workflow, and application pulling from a different, ungoverned understanding of the business doesn't produce intelligence; it produces inconsistency at scale. Shared context is becoming as fundamental to the AI stack as the database was to the application stack." The same shift is visible across the broader AI ecosystem. As Remy Thellier, Head of AI/ML Partners at Snowflake, notes, "C-suites aren't asking for another AI pilot. They want to see it working in production, with numbers behind it. Our joint customers get there by keeping AI analysis grounded in governed data on Snowflake, and now Snowflake's AI functions run right inside the Hex workflow. That's the difference between a demo and something that can actually be leveraged every day to drive significant value for the organisation." AI doesn't have a model problem. Many enterprises believe they are AI-ready because they have invested in cloud platforms, data warehouses and business intelligence tools. Yet, these architectures were built for analytics and reporting, not for autonomous AI agents capable of continuously reasoning across workflows. Different departments often calculate the same business metric differently. AI assistants generate conflicting recommendations because they interpret business logic independently. Employees spend more time validating AI-generated outputs than acting on them. According to Hex, these are not failures of AI models but of fragmented context. Hex argues that many organisations delay AI adoption while trying to perfect their data estate. Instead, those seeing measurable returns begin with a recurring business process, assemble enough trusted context to support that workflow, and allow their data foundation to mature as AI delivers value. Rather than treating data quality as a prerequisite, they improve it by solving real business problems. Shared context as AI's competitive advantage. For Hex, the defining challenge of enterprise AI is not connecting more data. It is ensuring every analyst, application and AI agent interprets data in exactly the same way. Business metrics such as revenue, customer lifetime value or churn often have different definitions across departments. Human analysts understand these nuances through experience, but AI agents cannot. Two systems can access the same dataset yet arrive at different conclusions if they rely on different business definitions. Aguilar believes the semantic layer is becoming a critical component of the AI stack. Rather than embedding business logic into individual dashboards or applications, organisations should create reusable semantics that every analyst and AI agent can build upon. Shared business definitions, governed metrics and common analytical logic become organisational assets rather than project-specific configurations. Their customer experiences reinforce this philosophy. At ClickUp, combining product usage, customer engagement and billing data enabled AI to identify high-value churn risks, helping customer success and marketing teams personalise interventions and save more than $1 million in revenue churn. At Huckberry, bringing together historical sales data, promotional calendars and finance expertise produced more accurate inventory forecasts and savings exceeding $1 million annually. In both cases, AI generated value because it had access to shared business context rather than isolated datasets. For the company, AI should complement analysts rather than replace them. Analysts contribute business judgement and institutional knowledge, while AI contributes speed, exploration and automation. Better decisions emerge only when both operate from the same trusted context. Open infrastructure is not just about moving data. As organisations prepare for an increasingly agentic future, open data infrastructure is becoming a strategic necessity. Open table formats such as Apache Iceberg, interoperable APIs and portable architectures allow enterprises to move data across technologies without repeatedly rebuilding pipelines. But Hex argues that openness cannot stop with storage. As organisations adopt multiple foundation models and specialised AI agents, the challenge is no longer moving data between systems, it is preserving business meaning wherever that data goes. An Open Data Infrastructure therefore extends beyond open formats to include an open semantic layer, ensuring business definitions, governance policies and analytical context remain portable across tools. Without context portability, data portability alone simply recreates fragmentation. Rather than waiting for every data quality issue to be resolved, Hex advocates starting with a single, measurable business process, building enough trusted context to make AI useful, and allowing the foundation to mature through production. Every successful workflow expands the organisation's context layer, making future AI initiatives easier to deploy. Building for Measurable AI ROI. Hex's CIO playbook encourages organisations to begin with business outcomes rather than platforms, identify where trusted context already exists, design workflows that business teams can own, measure success through operational outcomes rather than AI metrics, and allow the data foundation to evolve alongside real work. This philosophy is evident across Hex's customer stories. At EliseAI, AI reduced quarterly business review preparation from days to minutes while keeping analysts in control of the underlying reasoning. At Supabase, AI consolidated information spread across eight systems into a single workflow, enabling support engineers to resolve customer requests in seconds while maintaining human oversight before actions were executed. In both cases, AI succeeded because it became part of an existing business process rather than another standalone application. Boardrooms will increasingly judge AI investments by business outcomes rather than technical achievements. Organisations that start with real business processes, build just enough trusted context to solve them, and allow that foundation to mature through production will be best positioned to realise measurable AI ROI. Your reaction Discussion. No comments yet - be the first to share your view. Scoops & predictions tracker
dbt in 2026: core, Cloud, and the Mesh decision. Data Platforms Architecture 1. Problem. Every data platform that reaches multiple teams hits the same wall: the transformation layer grows into an unmaintainable monolith. Models are duplicated, ownership blurs, and one change in a shared table breaks three dashboards that no one can name. The classic failure is the "big ball of models": a single dbt project with hundreds of SQL models, one team owning the repo, and every other team waiting in a queue. The 2026 State of Analytics Engineering report, the largest industry survey from dbt Labs, measured the underlying shift: trust in data and data teams rose to 83% priority and speed of delivery to 71%, both accelerating faster than any other objective, while cost stability is now in the top three for 52%. This is the context where the real decision happens. Not "which SQL framework", but where do you run the transform layer, and when do you split it. 2. Context. The tooling is now two products with a clear split. dbt Core is the open-source transform engine: models, tests, sources, and the dbt run CLI that turns SQL and data tests into a query DAG. dbt Cloud adds a managed editor, scheduler, observability, the Semantic Layer that exposes metrics, and the "dbt Wizard" agent released in 2026 that proposes model content and runs checks in the CLI loop. Core supports the same model syntax and is free, but requires you to operate observability, scheduling, and CI/CD yourself. The practical boundary is sizing. Pragmatic field reports from 2026 consulting practices converge: teams below roughly 15 analytics engineers are better off running dbt Core plus an existing orchestrator (Airflow, Dagster, Prefect) because the operational cost of Core is small and the existing scheduler already runs their jobs. Above roughly 15 engineers, Cloud's managed scheduler, observability and Semantic Layer justify the per-seat cost. dbt Mesh, which fuses multiple dbt projects into a cross-project DAG with explicit contracts (YAML model contracts, versioned model groups, cross-project references), is only worth it above ~25-30 engineers split across 2-3 teams; below that production Mesh is overhead. 3. Trade-offs. | Option | Control | Governance ceiling | Cost model | CI/CD effort | | dbt Core + Airflow/Dagster | Full - runs on your infra | Single project, one repo | Free (runtime infra only) | Jacob orchestrator already does it | | dbt Cloud | Vendor-runtime | Single project, managed observability + Semantic Layer | Per-seat subscription | Built-in | | dbt Mesh (Core or Cloud) | Full + contracts | Multi-project, cross-team contracts | Added complexity + (Cloud seats) | Increases with cross-project tests | 4. Decision. For a self-hosted engineering team, the 2026 recommendation in three steps: * Start with dbt Core on your existing orchestrator. Until you approach 15 analytics engineers, the infrastructure already runs jobs; Core adds models and tests without adding a second control plane. * Adopt quality gates in CI as a first step, not an afterthought. The report's finding - trust (83%) and speed (71%) are the top priorities for 2026, above cost (48%) - is the roadmap: define model tests (dbt test), source freshness, contract tests on public models, and a -fail-fast CI gate. * Involve Mesh only when the DAG boundaries appear. When two teams start editing one project and code review turns into arbiter of a global namespace, that is when you split the project and publish access contracts via Mesh. Below that, one project with selectors and vars is simpler than the full federation. Contract-as-code. The structure that stabilizes the Mesh decision is the model contract. version: 2 contracts declare required columns and types; dbt build enforces them at compile time, not just in tests. This is the same discipline the design-firm uses for cross-team APIs. Before Mesh, make contracts mandatory on any model marked published. The runtime split. Keep the orchestrator as the sole scheduler; dbt stays a step in the DAG. Airflow/Dagster/Prefect trigger dbt run -selector production and handle retries and SLAs. You avoid the Cloud scheduler dependency at the 15-15 matrix only when you want the Semantic Layer or managed observability - and the Wizard shows the AI pair-programming direction the vendor is pushing, which matters for the platform's future, not the current core. 5. Result. With Core + CI + contracts, a 12-engineer team runs 3-6k rows of tests, with zero Cloud seat, in a DAG that still fits a single project. When the second team starts and change conflicts rise, the -contract + Mesh split lands in a weekend of mechanical work - because the disciplined core already separated models into layers and marked the shared ones public. The resources list stays small: models are SQL tests in YAML, runtime is a container. The report's 2026 signal - teams shipping faster via AI-assisted codegen while governance lags - matches the model itself: dbt is the winner of the "code correctness" workflow because tests and contracts are the governance. If you're designing the data layer from scratch, read the dlt loading pipeline post for the ingestion side of the same stack, or the sovereign warehouse case study for the full platform shape. Related reading. Have a specific infrastructure problem? Let's talk.
Northgrain Data joins the dbt Labs Partner Program. Northgrain Data has joined the dbt Labs Partner Program, formalising its focus on production-ready dbt implementations, migrations and analytics engineering. Northgrain Data has joined the dbt Labs Partner Program as a Registered Consulting & Services Partner. This is less a change in direction than a formalisation of how Northgrain Data has been building data platforms from the beginning. dbt sits at the centre of most analytics engineering work Northgrain Data deliver. Northgrain Data use it to turn fragmented transformation logic into maintainable data models that can be reviewed, tested, documented and deployed as code. Joining the partner program allows Northgrain Data to deepen that focus while working more closely with the dbt ecosystem. Why Northgrain Data build around dbt. Moving data into a warehouse is only the first part of building a useful data platform. The harder problem begins when an organisation needs to turn raw source data into shared definitions of revenue, customers, inventory, retention or operational performance. Without a structured transformation layer, that logic usually ends up distributed across: * spreadsheets * dashboard queries * stored procedures * standalone SQL scripts * Python jobs * individual analysts' knowledge The result is rarely one dramatic failure. It is a growing collection of smaller problems: duplicated logic, inconsistent metrics, undocumented dependencies and changes that are difficult to review safely. In its projects, dbt becomes the place where that business logic is structured and maintained. Models are kept in version control. Changes go through review. Tests run before incorrect data reaches downstream reports. Documentation and lineage make it easier to understand how a source field eventually becomes a business metric. How Northgrain Data use dbt in client projects. Its dbt work typically falls into several areas. Building new analytics platforms. Northgrain Data design and implement the transformation layer as part of a complete data platform, including ingestion, warehouse architecture, orchestration, testing, monitoring and deployment. Migrating existing reporting logic. Northgrain Data move transformations out of spreadsheets, BI tools, stored procedures and disconnected SQL scripts into a structured dbt project. This creates a single place where business logic can be reviewed, tested and maintained. Refactoring existing dbt projects. A dbt project can run successfully while still becoming increasingly difficult to work with. Northgrain Data review project structure, model dependencies, materialisations, naming, testing coverage, documentation, CI/CD and warehouse performance, then implement the changes required to make the project easier to operate. Productionising analytics workflows. Northgrain Data introduce the engineering practices required to move from a collection of models to a production system: * automated testing * CI/CD * deployment environments * source freshness checks * documentation and ownership * monitoring and alerting * operational runbooks What joining the partner program changes. The partner program gives Northgrain Data access to additional enablement, ecosystem resources and opportunities to work more closely with dbt Labs and other organisations in the dbt ecosystem. For clients, its delivery model remains the same. Projects are implemented in the client's environment. Code stays in the client's repositories. Testing, documentation, deployment and knowledge transfer are included in the scope rather than treated as optional additions. The partnership strengthens the technical direction behind that delivery model. Working with Northgrain Data. Northgrain Data currently support dbt and data platform initiatives through three engagement models. Data Platform Audit. A structured review of an existing data platform or dbt project, followed by prioritised findings and a practical implementation plan. Data engineering sprint. A fixed-scope implementation delivered over approximately three to six weeks. Typical projects include dbt migrations, platform builds, reporting automation and existing project remediation. Embedded data engineering. Ongoing dbt, analytics engineering and data platform support for teams that require additional delivery capacity. To discuss an existing dbt project or planned implementation, start with a Data Platform Audit.
Data contracts: the discipline that stops bad data before it becomes your problem. Every organization that moves data between systems has experienced the same failure: something breaks downstream, and by the time anyone notices, the problem has already propagated through multiple processes. You trace it back and find that a source system changed its structure without warning - a column removed, a format altered, a field shifted - and your pipeline absorbed that change silently until the damage was done. Data contracts exist to prevent exactly that, and they are beginning to gain the traction in the industry that they deserve. What a data contract actually is. The name is intentionally literal. A data contract is a formal agreement between the sender and receiver of data - the source system and the target system - that defines exactly what the data will look like. The structure, the format, the attributes present, the data types expected. Both parties commit to that specification: the source guarantees it will supply data in the agreed form, and the target confirms it will accept and process data in that form. If the agreement is broken, that is a violation - and it is treated as such. The critical distinction from a conventional agreement is that a data contract is not a piece of paper that people can quietly ignore. It is electronically enforced. The target system checks incoming data against the contract before loading it. If the data violates the specification, it does not get loaded. The violation is flagged within the governance structure, and the source system owner is alerted to fix it. The pipeline may pause, but bad data does not propagate. Preventative vs. Reactive: why the distinction matters. Most data quality management is reactive. Data is loaded, checks are run afterwards, problems are discovered, and teams work backwards to unpick the damage. This approach is expensive and slow, and it means that by the time a quality issue is identified, it has often already influenced reports, decisions, or downstream processes that relied on the data. Data contracts shift this to a preventative model. The check happens before the data enters the target system. This is a meaningful operational difference. Rather than discovering three days after a load that a column shift has put every field under the wrong name across thousands of records, you catch the violation at the point of entry. The source is notified immediately. Nothing downstream is contaminated. There is a useful analogy in lean manufacturing. When something goes wrong on a production line, you pull the handle and stop the line immediately. You fix the problem at source rather than allowing defective output to continue accumulating. Data contracts apply exactly that logic to data pipelines. This is not an entirely new idea. It is worth being clear that data contracts are not a radical departure from existing practice. Teams have always run schema checks on incoming data. The world of APIs has operated with a version of this concept for years - when you declare an API version, you guarantee that consumers calling that version will receive a specific, predictable response. Call a later version and the contract may differ, but the consumer understands and chooses that. Data contracts extend this discipline more formally and more broadly across the data exchange landscape, applying it not just to API interactions but to system-to-system data flows of all kinds. What is relatively new is the concrete tooling to implement them. DBT has released an implementation of data contracts that Business Thinking Limited has been using consistently in its data load patterns, and it works well. The field is still developing - implementations vary, and teams are finding their way to approaches that work reliably at scale - but the direction of travel is clear. Who should be thinking about this now. If data contracts sound extreme for where your organization currently sits, that reaction is informative. It likely means you are early in your data governance journey, and the overhead of implementing contracts formally may not yet be justified by the complexity of your data landscape. That is a legitimate position for now. But if your organization is actively developing data products, dealing with recurring data quality issues, or beginning to take data governance seriously as a discipline, data contracts are the right next step. They give you control over the flow of data in a way that informal agreements and after-the-fact checks simply cannot. And if you are at the point of designing new APIs or new data exchange mechanisms, that is precisely the right moment to establish contract agreements - while the structure is still being defined rather than retrofitted later. The concept was largely formalized by Andrew Jones, whose work on data contracts is worth reading for any team looking to implement them seriously. Chad Sanderson has also written substantively on the topic. There is a growing body of literature and practical guidance available, and the tooling is maturing quickly. The broader principle. Underneath the technical detail, data contracts represent a governance posture: the belief that it is better to stop bad data at the boundary than to absorb it and deal with the consequences. That posture requires agreement and accountability on both sides of every data exchange, which in turn requires organizational alignment that goes beyond the data engineering team. For senior leaders, the implication is straightforward. Data contracts are not just an engineering pattern. They are a governance commitment - a decision that data quality is enforced proactively, that source system owners are accountable for what they send, and that the integrity of downstream systems is protected by design rather than by luck. Ready to take the next step? If this has raised questions about how your organization manages data quality and governance across your platform, there are three ways Business Thinking Limited can help. Talk to Business Thinking Limited. If you want a direct conversation about your data architecture and where the pressure points are, get in touch. Business Thinking Limited work with senior data leaders and C-suite executives to cut through complexity and build platforms that perform. Download its white papers. Business Thinking Limited publish in-depth guidance on Data Vault, Medallion architecture, AI readiness, and modern data platform design - written for practitioners and leaders alike. Take the data transformation readiness assessment. Not sure where your platform stands? Its free diagnostic evaluates your organisation across four critical dimensions - platform foundation, reporting and insight speed, operational agility, and strategic positioning - and gives you a personalised report with a clear picture of where to focus. It takes less than ten minutes.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
Company Size
1,001-5,000
Company Stage
Series D
Total Funding
$414.4M
Headquarters
Philadelphia, Pennsylvania
Founded
2016
Find jobs on Simplify and start your career today