Jobtie/Jobs/Site Reliability Engineer

Site Reliability Engineer
- New York, NY, US +2 more
- Remote
- Full-time
- $140K/yr - $200K/yr
About the role
Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.
About the Role
- Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production.
- Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage.
- Design the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production.
- Own firmware-level telemetry, log collection at scale, and the low-level access layer that repair automation and health tooling depend on.
Skills & Experience
- You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it
- You're fluent with AI tooling. You aren’t afraid to max-out your token usage for the right spec.
- You’re comfortable debugging production issues, from triage to post-mortem.
- Enthusiasm for developer tools, cloud native technologies, and open source software
Benefits
- Competitive salary and meaningful equity
- Join a fast-growing pre-series A company at the ground floor
- Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
- Opportunities to participate in events across the cloud native community
- Fitness stipend, learning budget, and much, much more
About the company
Cloud computing is broken.
AI has introduced a new generation of workloads, like GPU inference, sandboxes, and agents. These aren't ordinary applications that can be run as Lambdas, or Dockerized apps on VMs: they're massive, stateless containers that need to spin up in <1s, often across multiple clouds and regions.
Today, engineers are hacking together infra that breaks under real-world loads. That's where we come in.
Our mission is to build the world's best compute platform for AI. Our first product is a serverless inference platform, used by companies like Coca Cola, Geospy and hundreds more. We've built our own container runtime, called beta9, which is designed for launching GPU-backed containers in under 1s.
We're a small, highly-technical team, with backgrounds in distributed systems and robotics. We've raised $7M from YC, Tiger, Guy Podjarny (Founder of Snyk), and Jason Warner (former CTO of Github).
We're searching for intensely curious, passionate, and hard-working engineers to join our mission in rebuilding the cloud for the age of AI.