All roles

Principal Data Platform Engineer

Full-time•Office-first (Bangalore)•Engineering

We are building autonomous infrastructure, and it runs on telemetry data.

base14 builds Scout, an observability platform built on OpenTelemetry that covers logs, metrics, traces, APM and RUM. Every signal our customers send lands in our telemetry data platform. When a customer's production is failing, Scout is how they see what is happening, so our platform has to be up when theirs is down.

We are also building agents that use this telemetry, both detailed and summarised, to manage production infrastructure. Self-healing, autonomous infrastructure is a company goal, and those agents are only as good as the data platform underneath them.

As we onboard more customers, we are investing in systems that hold 99.99% uptime and sub-second ingestion for gigabytes of telemetry data. We are hiring a Principal Data Platform Engineer to own that platform and set its technical direction.

The Role

You will own the telemetry data platform end to end. Data arrives through OpenTelemetry Collectors, Kafka, or Parquet written directly to S3, depending on the scale and type of pipeline. Raw data lives in ClickHouse. Summary data flows from there into OLTP stores, graph databases and caches that serve different parts of the product.

ClickHouse will take a good share of your time: running clusters, designing them for high availability, and keeping queries fast as volume grows. The platform is wider than one store, though. It runs from the ingestion pipelines through to the summary stores and caches that serve the product, and you own how those pieces fit together. This is an engineering role as much as an operations role. You will write the services, tooling and automation around the data as well as run it.

As a principal engineer, you make the architectural calls for the data platform, write them down, and raise the bar for the engineers around you.

You will share the on-call rotation for the data platform. We expect you to treat every page as a defect and engineer it out of the system.

We expect you to automate your own workflow. You will use AI tools and LLMs to write code, generate tests, analyze slow queries, draft migrations and runbooks, and investigate incidents, so you can focus on hard technical problems.

What You'll Do

  • Run ClickHouse in production: Operate clusters across AWS, GCP, Azure and our data centers. Own sharding, replication, capacity planning, upgrades, backups and recovery.
  • Engineer for four nines: Design the high-availability architecture for the data platform. Plan for node, zone and region failures, test those plans, and remove single points of failure.
  • Keep queries fast: Design schemas, sorting keys, partitioning, materialized views and TTL policies for logs, metrics and traces. Find and fix slow queries before customers notice them.
  • Build ingestion that keeps up: Evolve our pipelines across OTel Collectors, Kafka and Parquet on S3 so ingestion stays sub-second as volume and customer count grow. Handle bursts and backpressure without losing data.
  • Write the services around the data: Build the services that write to and read from ClickHouse, and the jobs that move summary data into PostgreSQL, Neo4j and Redis.
  • Make the data usable by agents: Shape detailed and summary telemetry so our agents can query it quickly and trust what they get back when they act on production infrastructure.
  • Automate the infrastructure: All of our infrastructure is automated and delivered through GitOps with Argo CD. Every infrastructure change you make ships as a reviewed commit.
  • Share on-call: Take part in the shared rotation for the data platform. Lead incident response when the data layer is involved, write the postmortem, and fix the root cause. Automate the Incident response for areas owned.
  • Set technical direction: Decide how the data platform evolves. Review designs, write down the tradeoffs, and mentor engineers working on the data path.
  • Automate your work: Use AI assistants and automation tools to write, test and debug code, and to speed up operational work.

What We Look For

Must Have
  • Production ClickHouse at scale: You have run at least one ClickHouse cluster in production holding terabytes of data. You have dealt with replication issues, resharding, upgrades and node failures on a live system.
  • High-availability experience: You have designed and operated HA infrastructure for a database, ClickHouse at minimum. You know replication, failover, coordination with ClickHouse Keeper or ZooKeeper, and backup and restore from having done them.
  • A programmer first: You have built production services that use ClickHouse. Your experience goes well beyond writing queries and operating the database. You write clean, legible and fast code. Most of our data path is Go and Rust.
  • Schema and query depth: You understand the MergeTree engine family, how sorting keys and partitions affect reads and merges, and how to find a bottleneck using query logs and system tables.
  • Operational ownership: You are comfortable carrying on-call for the systems you build, and you prefer fixing root causes to restarting things.
  • AI-assisted execution: You use generative AI tools to write code faster and to handle routine engineering and operational tasks.
  • Principal-level judgment: You can take an ambiguous scaling problem, choose a direction, explain the tradeoffs in writing, and bring other engineers along.

Years matter less to us than what you have run. People at this level typically have 10 or more years of engineering experience, but depth with ClickHouse in production counts for more than the number.

Strong Pluses
  • Agent-building experience: You have built AI agents or LLM-driven systems that take real actions, and you are excited about infrastructure that heals itself. This is a massive bonus for us.
  • Kubernetes: You have run stateful workloads on Kubernetes, including operators, persistent storage and rolling upgrades of databases. This is a big advantage here.
  • Observability background: You have worked on logs, metrics or traces at scale, or with OpenTelemetry, and you understand the shape of telemetry data.
  • More scale: You have managed multiple ClickHouse clusters, or clusters holding petabytes.

Technical Stack

  • Analytics & Storage: ClickHouse, Apache Pinot, S3, Parquet
  • Ingestion: OpenTelemetry Collector, Kafka, Parquet direct to S3
  • Summary & Serving: PostgreSQL, Neo4j, Redis, both self-managed and cloud-managed
  • Languages: Go and Rust, with Python, Ruby and TypeScript where needed
  • Infrastructure: Managed Kubernetes on AWS, GCP and Azure. Talos Linux on VMs in our data centers.
  • Delivery: Argo CD, GitOps, 100% automated infrastructure
  • AI & Automation: AI coding assistants, LLM APIs and automation tooling

Why Join Us

  • Build autonomous infrastructure: We are building agents that read telemetry and manage production systems on their own. The platform you own is what they see and reason with, so your work decides how far self-healing infrastructure can go.
  • Hard problems at real scale: Four nines of uptime and sub-second ingestion on a multi-cloud telemetry platform is a problem few engineers get to own.
  • Real ownership: You set the direction for the data platform and your decisions shape the product and the company.
  • Equity: Every employee owns a stake in the company.
  • Strong peers: You will work alongside engineers who have scaled massive infrastructure.
  • Fast feedback: Your work shows up in customer-facing latency and uptime within days.

What We Offer

  • Competitive salary and equity.
  • Collaboration with experienced founders.
  • Hardware and software of your choice, including premium AI development tools.
  • Learning budgets and flexible hours focused on work output.

How to Apply

Email hello@base14.io with your resume. Tell us about a ClickHouse cluster you ran: how big it was, what broke, and what you changed because of it.