Do you only do ops, or also design and build?
Both. We design and build data platforms and pipelines, ML and GenAI systems, and analytics foundations — then we own the ops so they stay cheap, reliable and accountable. Most engagements start with build or restructure; the teardown maps which you need first.
Do you replace our existing stack?
No, not by default. If you run Databricks, Snowflake, BigQuery or Kafka we work inside it. We propose a base-layer change only when the current one is the actual cause of the problem, and we will show you the arithmetic before you decide.
When does DataKru make sense instead of a lakehouse?
When you have no platform engineers, data volumes in the gigabytes rather than petabytes, and you want a lean base without a lakehouse contract. DataKru itself is in free beta (pipelines and on-demand jobs today); scheduling, observability and governance arrive via the product roadmap or via our services engagement. Above that scale, or where a platform team already exists, keeping your stack is the better call.
What actually happens in the first 30 days?
Week 0 is the free teardown. Weeks 1–2 are the architecture sprint: we map sources, transforms and consumers, run a gap analysis and deliver a written report plus a 90-day backlog with a live readout. Weeks 3–4 are either your team executing that backlog, or a scoped implementation SOW if you want us hands-on.
How does pricing work?
The teardown is free. The architecture sprint is $2,000 fixed and is the only price on this page for a reason — it is the one thing we can scope without knowing you. Implementation is a separate SOW priced from the sprint backlog. Retainers are monthly.
What does LLMOps mean in practice here?
We design and ship the GenAI system first when it does not exist — then the ops layer: an evaluation set before a prompt or index change ships, versioning on prompts and retrieval config, cost per answer as a first-class metric, and a human gate on anything the model should not decide alone. We will also say when something should not be AI and help you remove it.
Is AIOps just alerting with an LLM bolted on?
No. The order matters: design the observability and cut alert noise first, then automate diagnosis against runbooks, then automate remediation — and remediation stays behind a human gate until the diagnosis has earned trust. Skipping the first step gives you a confident model reasoning about noise.
Where does this run — your servers or ours?
Either, and it is a separate decision from whose stack you keep. We can deploy into your own AWS, GCP, Azure or on-prem environment, where data and credentials never leave your boundary and your team holds the keys. Or we host and operate it on our infrastructure under a retainer, which is usually the right call when there is no cloud or platform team. Because we build on open formats and standard components, starting hosted and moving in-house later is a planned migration rather than a rebuild — and we document the exit as part of handover.
If you host it, who is responsible for the data?
You remain the data owner; we are the operator. We will sign an agreement covering access, retention and deletion before anything of yours lands on our infrastructure. Being straight about the limit: we are an engineering team that runs infrastructure, not a certified hosting provider. If your procurement requires certified infrastructure or a data-processing arrangement we cannot meet, we deploy into your account instead — and we would rather tell you that in the teardown than at contract stage.
Do you need production access to our data?
Not for the teardown or most of the sprint. Schemas, DAG definitions, sample volumes and a conversation with whoever gets paged are usually enough. Implementation access is scoped and least-privilege under your controls.
What is DataKru Atlas, and can we buy it?
Atlas is a planned context-graph product for finding high-intent buyers, starting with auto and mobility. It is roadmap and early conversations only. We will not sell it into a delivery contract until it ships.