Baseten•5h ago
Career Pages
Software Engineer
San Francisco, CA
Full Time
Senior Level
Full Job Description
Role Overview
Baseten powers mission-critical inference for dynamic AI companies by combining applied AI research, flexible infrastructure, and developer tooling to bring advanced models into production. The Software Engineer will build and operate distributed systems for large-scale LLM inference.
Key Responsibilities:
- Build the infrastructure and orchestration systems that deploy and run large-scale distributed LLLM inference, including routing, autoscaling, scheduling, and runtime management
- Design, build, and operate Model APIs with a focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling, and multimodal serving
- Implement platform fundamentals such as API versioning, validation, usage metering, quotas, and authentication
- Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks for speed, reliability, and quality. Help set best practices for testing, release automation, and operational excellence
- Debug and harden complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads to improve reliability and scalability
- Partner with Inference Performance engineers and other teams to make new optimizations broadly available to customers and easy to configure
- Own projects end-to-end from architecture through deployment, monitoring, and iteration on customer feedback. Make thoughtful tradeoffs between performance, reliability, operational simplicity, and developer experience
Company
Baseten
Baseten provides essential infrastructure, specialized tooling, and expert support to seamlessly integrate artificial intelligence into business operations.
San Francisco, CA
Posted on Career Pages