LakehouseLakehouseTalent Ecosystem

Data engineering — a Lakehouse ecosystem

Your Lakehouse Works. Your Capacity Doesn't

Riley Spraggs ·

A platform evaluation ends with a signed contract, a reference architecture and a roadmap. Nine months later the roadmap has slipped, the warehouse bill has not, and the post-mortem reaches for the platform: wrong engine, wrong table format, wrong partner. It usually isn’t. The Unity Catalog works. The Iceberg tables land. Streaming ingestion runs. What broke is the number of people who can turn a backlog of use cases into production pipelines at the pace the business was promised — and that is a staffing problem wearing a technology problem’s clothes.

The barrier is implementation capability, not the technology

When data teams are asked directly what stops them adopting open table formats, the answers are not about performance, cost or lock-in. The top barriers are knowledge gaps at 27% and unclear use cases at 27% — teams navigating implementation complexity rather than resisting the concept. Read that as a hiring signal. Nobody is saying Delta or Iceberg can’t do the job. They’re saying they don’t have enough people who have done the job before, and no agreed list of what to point them at first.

The same pattern shows up on the AI side of the house. Gartner found that organizations reporting successful AI initiatives invest up to four times more as a percentage of revenue in foundational areas — data quality, governance, AI-ready people and change management. Three of those four line items are people and process. The platform is the cheap part of the gap.

Nobody ever filed a ticket that said “blocked: insufficient senior engineers.” They file tickets that say “blocked: waiting on data.”

How to tell a capacity problem from a platform problem

The two failure modes look identical from a steering committee. They are easy to separate if you ask the right question.

Symptom Platform problem Delivery capacity problem
Pipelines fail Same failure, reproducible, vendor-acknowledged Different failure each time, usually a pattern nobody standardized
Roadmap slips Blocked on a feature on the vendor’s roadmap Blocked on one named person’s calendar
Cost overruns Inefficient engine or pricing model Nobody owns cluster policies, warehouse sizing or job tuning
Governance gaps Tooling genuinely can’t express the policy Tooling can, nobody has had time to implement it
New use cases stall Needs unsupported workload type Queue behind existing migration work
Adoption flat Consumers can’t get what they need from the tool Semantic layer and documentation never got built

The tell in the right-hand column is that every row resolves to a person, not a product. If your top three blockers all end in a name, you have bought the right platform and under-resourced the team operating it. We walk through that diagnostic in more detail in capacity versus adoption.

The partner-side version of the same test

If you sell lakehouse implementation, you see this from the other end. Consumption targets miss not because the customer churned but because the customer could not absorb delivery — their engineers were seconded to a migration, their data stewards never got appointed, their one person who understood the source systems left. A stalled account where the platform works fine is almost always a customer-side capacity account, and it responds to staffing augmentation, not to another architecture review.

Spend is growing faster than the team

The clearest quantitative signature of a capacity problem is spend outrunning headcount. 57% of data teams report increased warehouse and compute spend, against only 13% reporting decreases. Nobody reports their team growing at that rate. So compute per engineer climbs, and the engineers you have spend an increasing share of their week on operations, cost triage and firefighting rather than on the use cases that justified the spend in the first place.

Meanwhile the pressure to ship got worse. Speed of delivery as a stated priority jumped from 50% to 71% year over year. Same team, more compute, sharply higher expectation of velocity. That is the arithmetic of a roadmap slipping while every individual on it works harder than last year.

And the payoff is not landing. Only 39% of technology leaders are confident their enterprise’s current AI investments will have a positive impact on financial performance. A 39% confidence rate across a cohort that has already spent the money is not a tooling verdict. It is a conversion verdict — spend went in, delivery didn’t come out the other side at the rate anyone modelled.

Maturity beats features

The instinct when a roadmap stalls is to buy the next thing: a catalog, an orchestration layer, an agent framework. The evidence points the other way. Organizations with the highest maturity of AI-ready data and analytics capabilities are achieving up to 65% greater business outcomes, measured in revenue growth and cost optimization. Maturity is not a SKU. It is accumulated practice — someone who has already made the lineage decisions, already written the cluster policies, already argued the medallion boundaries with a business stakeholder and lost once.

That is why the specialist profile matters more than the headcount number. A platform engineer who has run Unity Catalog migrations at two prior companies removes six weeks of discovery that a generalist has to live through. Same cost line, different slope on the curve.

What the unblocking roles actually are

Blocker Role that clears it Typical engagement shape
Migration backlog from legacy warehouse Senior data/platform engineer with prior migration reps Fixed-scope, 3–9 months
Cost per query climbing Lakehouse performance specialist Short, surgical, often part-time
Governance policy unimplemented Unity Catalog / governance engineer Project, then handover
Consumers can’t self-serve Analytics engineer owning semantic layer Permanent in-house
Streaming use case parked Streaming specialist (Structured Streaming, Kafka, Flink) Project with in-house shadow
Nobody owns the standards Staff-level lead Permanent, hire first

The column that gets skipped is the third one. Not every blocker wants a full-time hire, and not every blocker should be outsourced. Getting the shape right is most of the saving.

Why you can’t simply hire your way out at speed

The labour market headlines are misleading here. The broad US hiring market cooled through 2025 — Indeed’s Job Postings Index began the year more than 10% above pre-pandemic norms and slid to barely above them by late October. That reads like a buyer’s market, and for generalist roles it is. It tells you almost nothing about whether you can hire someone who has shipped production Delta Live Tables.

Specialist supply is moving the other way. US employment of data scientists is projected to grow 35% from 2025 to 2035, much faster than the average for all occupations. General slack and specialist scarcity coexist comfortably. So the practical constraint on a stalled roadmap is not the salary you’re willing to pay — it’s the eight to sixteen weeks of search, notice period and ramp before the person you hired is net-positive, which is covered in more detail in the delivery capacity gap.

A roadmap that slipped two quarters ago cannot wait another two for a permanent hire to ramp. That is the entire argument for blending.

The blended unblock: in-house core, flexible capacity around it

The pattern that works is not “hire a team” and not “hand it to a GSI.” It is a small permanent core that owns standards and institutional knowledge, surrounded by flexible specialist capacity sized to the backlog at any given moment.

Hire in-house for: platform ownership, cluster and cost policy, the semantic layer, stakeholder relationships, anything where the knowledge needs to still be in the building in three years.

Buy flexible capacity for: migrations with a defined end, performance work, a parked streaming or ML use case, and anything where the skill is needed intensely for one quarter and never again. The GSI or in-house decision usually resolves on exactly that axis — duration and whether the knowledge needs to persist.

Two guardrails make the blend work rather than create a second mess:

  1. Every external workstream has an in-house shadow. One named person on your side who reviews the PRs. Without it you are buying delivery and renting the understanding of it.
  2. Standards are set by the core before capacity arrives. Naming, layering, testing, deployment. Five specialists working to three different conventions is slower than two working to one.

For partners, the same structure solves the absorption problem. When a customer’s capacity is the blocker, placing specialists into the customer’s team — not adding consultants to your own statement of work — restores their ability to consume what you’re delivering. Accounts un-stall because the customer can ship again.

A two-week diagnostic

Before you renegotiate anything with a platform vendor, run this:

  • List the top ten blocked items. Write the actual blocker next to each. Count how many resolve to a person rather than a product. If it’s more than six, stop evaluating platforms.
  • Divide compute spend by engineering headcount, this year against last. If that ratio grew, your spend is outrunning your capacity and the gap is the thing to fund.
  • Mark each blocker permanent or finite. Permanent goes to a hire. Finite goes to flexible capacity. Do not let the urgent ones default into whichever is faster to procure.
  • Name the shadow for every finite workstream. No shadow, no engagement.
  • Check the ramp maths. If the blocked item has a deadline inside four months, a permanent hire will not meet it, and pretending otherwise is how a quarter gets lost.

Most roadmaps that look stalled are not stalled — they’re queued. The platform you bought is running fine. The queue in front of it is the budget line you haven’t funded yet, and it’s staffed with people who are, per the projections, going to get harder to find rather than easier. When you’re ready to size it, the data engineering network is searchable without an account, and the Databricks hiring guide covers what the specialist profiles actually look like.

FAQ

How do I tell whether my lakehouse roadmap is blocked by the platform or by my team?

Look at what each blocked item resolves to. If your top blockers end in a person's name or a calendar rather than a vendor roadmap item, it is a capacity problem. When data teams are asked directly, the top barriers to adopting open table formats are knowledge gaps at 27% and unclear use cases at 27% — implementation capability, not the technology.

What is the fastest signal that platform spend has outrun delivery capacity?

Divide your compute spend by engineering headcount this year and last year. If that ratio grew, your platform spend is outrunning your ability to ship against it. Across data teams, 57 percent report increased warehouse and compute spend while only 13 percent report decreases, and almost none report headcount growing at that rate.

The hiring market cooled — why is specialist lakehouse talent still hard to find?

Because the two markets are different. The broad US job postings index slid from more than 10 percent above pre-pandemic norms to barely above them through 2025, while US data scientist employment is projected to grow 35 percent from 2025 to 2035. General labour-market slack does not translate into available specialist lakehouse delivery capacity.

What should be hired in-house versus bought as flexible capacity?

Hire in-house for platform ownership, cost and cluster policy, the semantic layer and anything whose knowledge must still be in the building in three years. Use flexible specialist capacity for finite work — migrations, performance tuning, a parked streaming or ML use case — with a named in-house shadow reviewing every external workstream.

Will buying another tool fix a stalled data roadmap?

Usually not. Organizations with successful AI initiatives invest up to four times more as a percentage of revenue in foundations including data quality, governance, AI-ready people and change management, and only 39 percent of technology leaders are confident current AI investments will improve financial performance. That is a conversion problem, not a tooling problem.

Databricks, Snowflake, Microsoft Fabric, Apache Spark and dbt are trademarks of their respective owners, used here only to describe specialists’ experience. Lakehouse is an independent talent network operated by Sloane Staffing; none of these vendors endorses or sponsors this site.

Size the delivery gap, not the platform

Search the data engineering network