Full-Time
Collaborative data analytics platform with notebooks
$160k - $245k/yr
San Francisco, CA, USA + 1 more
More locations: New York, NY, USA
Hybrid
Hybrid role; 2-3 days in office per week in SF or NYC.
See people who can refer or advise you
Hex.tech provides a collaborative data analytics workspace that unifies SQL, Python, R, and no-code in a single notebook interface, letting data teams analyze and visualize data without switching tools. It offers built-in AI that can generate queries, code, visualizations, and even start full analyses from a prompt, plus built-in data connections to warehouses, lakehouses, and databases. Real-time collaboration, version control, and reusable components support team workflows, with deep integrations to dbt, Snowpark for Snowflake, Spark, GitHub, GitLab, and workflow tools like Airflow, Dagster, and Prefect. The freemium model lets users start for free and pay for more features or capacity. Hex.tech’s goal is to make data analysis accessible and efficient for a wide range of users—from data scientists to business analysts and non-coders—by providing a user-friendly, end-to-end platform for querying, coding, visualizing, and sharing insights.
Company Size
201-500
Company Stage
Series C
Total Funding
$177M
Headquarters
San Francisco, California
Founded
2019
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Competitive salary & equity
Health, vision, & dental
Flexible PTO
Parental leave
401k matching up to 4%
Workstation + learning & development stipends
Hex releases DataBench: realistic agentic analytics benchmark. Tl;dr. Hex introduces DataBench, a 100-task benchmark for evaluating AI agents on real-world analytics work, revealing models struggle with judgment and avoid admitting uncertainty. Key points. * GPT-5.6 Luna dominates cost-efficiency, achieving near-Sol performance at 1/14th the cost on Pareto frontier * Claude Fable 5 only model avoiding performance regression at maximum effort levels; Opus 5 talks itself into wrong answers * Models score 75% on evidence-gathering Q&A tasks but only 54% on trap tasks requiring deeper reasoning and judgment * Frontier models rarely admit uncertainty; only Opus 5 passes collections-call-list trap requiring skepticism of plausible-but-wrong evidence Why it matters. Existing analytics benchmarks (Spider 2.0, DABstep) report 85-90% accuracy but test toy problems, not real work. DataBench exposes critical gaps: frontier models manufacture false confidence, struggle with ambiguous prompts, and regress when given more compute. This matters for anyone deploying AI agents on actual business analytics - you need skepticism and human oversight, not blind trust in high benchmark scores.
Credits and usage visibility for Hex Agents. Hex is rolling out monthly credit grants for Hex Agents, plus new ways to monitor usage and buy add-on credits. Jo Engreitz July 14, 2026 Since Hex launched the Hex Agent last fall, it's become the most capable analytics agent on the market. Every month, companies from startups to Fortune 500s trust the Hex Agent with millions of business-critical data tasks. Agents have actually become the primary mode of working in Hex, surpassing manual notebook cell creation and point-and-click exploration. Now, Hex is introducing a credit model that allows Hex to continue providing this experience, along with new ways to monitor usage and purchase add-on credits. How it works. Hex plans now come with a monthly credit grant for every paid user: * Professional Editor seats include 30 monthly credits * Team Editor seats include 40 monthly credits * Enterprise Editor seats include 60 monthly credits * Explorer seats include 10 monthly credits For more usage, Admins can purchase pooled add-on credits with auto top-ups that refill your balance as-needed. You can set a workspace-wide spend limit to stay in budget, and control who can use to add-on credits. You can even customize monthly add-on allocations per user at the workspace, group, or individual level. Change your mind later? No sweat, updates take effect immediately. Full details are in the docs. Existing customers will automatically move to this model with a grace period (annual contracts will need to be updated). New customers automatically start on this model. Effort-based consumption. Hex agents use credits based on effort - the complexity of the task, the amount of context the agent needs to process, and the resources required to complete it. This means simple questions cost very little, while more involved analyses cost more. A few illustrative examples: One big focus area for Hex is cost efficiency. Hex is constantly evaluating frontier and lower-cost models, choosing models that strike the best balance of accuracy and cost. Hex do this at a per-task level, too, using its subagent architecture. For example, some lower-cost models are already at parity with frontier models on search or simple data viz - so Hex direct traffic to those cheaper models, and reserve the expensive ones for jobs where their performance is worth the premium. The result: your credits go further over time, and you're not locked in to one model provider. You don't have to leave that entirely to Hex, either. The model picker lets you choose exactly which model powers your session, and how much effort the agent should exert: * Fable 5 and Opus 4.7 for your hardest, most open-ended analyses * GPT 5.5 as a comparable alternative to Opus * Sonnet 4.6 for everyday work * Kimi 2.7, for Opus-level performance at less than half the cost Not sure where to start? "Auto" is its out-of-the-box default, and it picks the best model for the task based on its own evals. Admins can also set a different model as the org-wide default, if that's a better fit for how your team works. For more on when to reach for each one, see Model Picker Best Practices. Full visibility. While this sort of effort-based model has become commonplace, Hex provides a unique degree of control to understand and influence your team's credit consumption. Users can view their credit balance at any time and monitor usage right down to the prompt level - click the three-dot menu on any completed agent task to see exactly what it cost. Admins get the wider view: a snapshot of every paid user's monthly credit balance, historical usage logs, add-on credit options, and workspace spend limit settings, all in one place. And Admins can go even deeper with credit usage visibility in the Context Studio - a great way to understand what topics users are relying on AI for, and to identify domain areas that could use more context curation. Managers and Admins can see what questions users are asking, which topics come up most, and where Hex Agents express uncertainty or raise warnings - all tied together neatly with proactive context suggestions. Data teams use that signal to tune context, update data definitions, and test changes before publishing. This drives better data accuracy for users, and better credit efficiency for agents. As organizations adopt AI more broadly, leaders are looking for better ways to understand the costs, so they can budget effectively and calculate ROI - and Hex is excited to give them ways to do that. Looking for some ways to optimize credit usage? Its Head of Data, Katie, wrote about tactics to ensure optimal credit efficiency here. This is something Hex think a lot about at Hex, where Hex is creating a platform that makes it easy to build and share interactive data products which can help teams be more impactful. If this is is interesting, click below to get started, or to check out opportunities to join its team.
Kimi K2.7 in Hex: near-frontier analytics at a fraction of the cost. Open models have finally caught up to their closed counterparts Matt Redmond Data teams July 8, 2026 Kimi K2.7 takes more turns than leading frontier models. It second-guesses itself, re-checks its own work, and usually finishes slower. So why does it cost less to answer the same questions - and get more of the hard ones right? Today Hex is excited to announce that Kimi K2.7 is available in Hex, hosted on US-based infrastructure! Hex believe this model performs at an intelligence level comparable to Opus 4.7, despite consuming 2.2x fewer credits on average. Hex find Kimi is empirically most effective on easier, non-visual questions. Reach for it on semantic-layer questions, BI questions, or any analytically hard work. Stick to Opus+ or GPT-5 class models for now on large, complicated visual understanding tasks. Kimi K2.7 is just the first piece of a deliberate, long-term bet Hex is making on open-weight models in general. Hex has been waiting for this inflection point for a long time, and have some really cool things coming soon! Why now? Data analytics has proven to be a weirdly difficult domain for agents. For the last year or so, whenever Hex has tested a Cool New Model in Hex that's not the latest big-lab frontier hotness, Hex has been disappointed. Overthinking, misunderstanding intent, unnecessary churning tool calls, poor instruction following, and other generally bad analytical behaviors were rife. Even models that performed competitively on public coding and math benchmarks were inadequate on its internal analytical benchmark. The messy context and vague, almost-but-not-quite verifiable tasks endemic to data workflows seem to be a perfect storm for LLMs - and particularly seem to throw smaller "benchmaxxed" reasoning models into paroxysms of overthinking, a failure mode exhibited not just by open-weight models but also by recent generations of Claude Sonnet. An open model can be as cheap and fast as it wants, but speed means nothing if it confidently hands you the wrong number when you ask how many turboencabulating widgets you made in Q3 last year. Hex feel that with Kimi K2.7, Hex finally have an open model that's on par with Opus and GPT-5.5 at data tasks. The open-weights tradeoff. Large language models in agent harnesses can be compared on three attributes: price per task, time per task, and success rate per task. These attributes directly map onto the concepts "economic efficiency", "system speed", and "model intelligence". There is no free lunch in this space - only tradeoffs. Would you buy a 20% speedup for 80% more cost? Would you sacrifice 5% model intelligence for answers that arrive 50% sooner? Anthropic and OpenAI offer a few guided choices in this tradeoff space - Fable is an "expensive, slow, but very smart" model intended for your hardest tasks, Haiku is the "cheap, quick, but less intelligent" model intended for easier ones. Sonnet and Opus exist in between. Open-weight models make available a completely new region of this tradeoff space. With its initial release of Kimi K2.7, you can choose lower credit usage at the expense of an increase in total end-to-end latency, without much of an intelligence haircut. And this is just for starters! Because Hex host this model ourselves, Hex can keep experimenting with new ways to configure these tradeoffs - this is the foundational promise of open-weight models. Hex could potentially offer a significantly faster Kimi at an increased cost, or a radically cheaper "batch" mode for slower async tasks. Hex chose Kimi for the native multimodality, but as other open models like GLM-5.2 approach the frontier, Hex can explore those too. The high cost of cheap tokens? It's tempting to reason about cost through the lens of model input/output token pricing, but in the agentic multi-turn world, that lens is too reductive. A brilliant model with "expensive" tokens can sometimes (often!) end up cheaper per task than a bargain-bin model with "cheap" tokens, simply because the smarter model is clever enough to get in, answer the question, and get out with fewer tokens and turns. Hex see this constantly. This paradox poses a problem for its open model initiative: its goal is to give you another point in the tradeoff space that's competitive with the frontier labs, but "cheaper per token" isn't sufficient. The model has to be smart enough to do the whole task cheaply - being cheaper per token doesn't matter if you ramble. And boy, can they ramble. Hex has watched cheaper models fumble their way through its harness, racking up expensive trajectories despite the ostensibly friendlier sticker prices on tokens. The most encouraging thing Hex can say here is that the gap is closing fast: as open-weight models get better, both their per-token pricing and total token usage have trended down, making them a more competitive option empirically in its testing. Hex is also co-evolving its harness with models over time, making the affordances more in-distribution with how the models want to use them. Hex haven't fully mitigated the rambling - Kimi K2.7 will sometimes yap more than Opus 4.7 - but it's a lot better than previous large open-weight models, and Hex believe it's crossed the threshold into "useful for most day-to-day analytics work" without being TOO annoyingly verbose. Most importantly, the per-token economics are sufficiently strong such that most tasks end up cheaper despite incurring more token usage. Evaluation results and discussion. Across its evals, some patterns emerge: Kimi is slower to finish tasks end-to-end, using more tokens and turns than Opus. However, due to favorable per-token economics, the price-per-task ends up substantially better in all domains except visualization, where it's only slightly better. The semantically modeled questions evaluation set contains tasks that a good semantic model can basically answer on its own - the bread-and-butter BI questions like "how many customers do we have". Here, Kimi and Opus do effectively the same. In fact, this set is effectively saturated now, and Hex'll be retiring it soon. This is one of the strengths of Kimi - quick, straightforward questions that you expect to be well modeled in your ecosystem will be inexpensive. The semantically unmodeled questions evaluation set contains tasks which cannot be answered by semantic models. It measures the agent's ability to "go off road" and explore the data warehouse directly. Opus is decisively better than Kimi at this task, winning by eight percentage points. The core visualization evaluation set measures how well the agent creates and interprets charts. This is the least favorable price comparison for Kimi - verifying data visually seems to be a difficult thing to do efficiently, so the turn count is larger. Hex is working on better harness support here, but for now this is the one domain where Kimi's economic efficiency is least impressive compared to Opus. The analytically hard evaluation set measures how well the agent is able to answer questions where the core complexity is in the analysis - even when context retrieval is easy. On this eval set, Kimi scores better than Opus, but takes about twice as long in the median case. This is due to an interesting emergent qualitative behavior: Kimi is a capital-W Worrier. K2.7 spends extra turns in validation and verification, second-guessing its own assumptions and re-checking its hypotheses. This pays off in correctness - it gets more questions right than Opus - but it also takes quite a bit longer. The contextually hard evaluation set measures how well the agent is able to answer questions when the analysis path is straightforward, but identifying the appropriate resources and context is difficult. On this set, Kimi performs about the same as Opus. Overall, Opus tasks take about 0.65x as long as Kimi tasks, but cost 2.22x as much. No free lunch. K2.7 has a smaller context window than most frontier models. Hex serve a 256k token context window for K2.7, but Anthropic and OpenAI's models have a 1M token context window. This means that threads will automatically compact more frequently on K2.7 than they would on a closed weight model, which means that some details from your working session may be compressed or forgotten earlier than they would otherwise. As mentioned, Hex also observed that while the average economic case is better with K2.7, there are still agent trajectories where K2.7 exhibits substantially longer reasoning behavior and incurs more agent turns than the closed weight models do. This appears in the long tail of the distributions - "most" Kimi agent runs are more efficient, but the rare ones that are worse are often quite a bit worse. What's next. Hex is excited to bring open-weight models to the Hex ecosystem, and Kimi K2.7 is the first experimental step rather than the finish line - it's unlocking deeper nodes on the tech tree for Hex. As always, Hex'll keep claims grounded in the data - when something gets faster, cheaper, or smarter, Hex'll show you the numbers. This is something Hex think a lot about at Hex, where Hex is creating a platform that makes it easy to build and share interactive data products which can help teams be more impactful. If this is is interesting, click below to get started, or to check out opportunities to join its team.
SAN FRANCISCO, CALIFORNIA / ACCESS Newswire / June 4, 2025 / Hex, the unified, AI-powered platform for analytics and data science, released its 2025 State of Data Teams Report today at Snowflake Summit 2025.
Hex, the first unified, AI-powered workspace for data science and analytics, announced a $70 million Series C funding round today led by Avra, with participa...