Staff AI Infra Engineer (serving API)

Company: Genmo Replay
Location: San Francisco
Posted on: November 6, 2024

Job Description:

We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.Role OverviewWe are looking for a senior/staff software engineer to join our inference team. In this role, you will be responsible for designing and scaling our inference systems as they grow to support over millions of users across more than 20 different data centers.Key Responsibilities

Develop high-performance, high-throughput, efficient, and low-latency inference pipelines.
Design, develop, and maintain scalable backend services that support our AI-powered content creation platform.
Implement and optimize model serving infrastructure using Kubernetes and other cloud-native technologies.
Collaborate with ML engineers to transition models from research to production.
Design APIs for integrating our AI capabilities into our partner ecosystem.
Implement monitoring, logging, and alerting systems for backend services and model inference.
Develop monitoring infrastructure for our ML serving pipeline and apply advanced model compression and optimization techniques (quantization, pruning, distillation) to improve inference performance.Qualifications
Bachelor's or Master's degree in Computer Science, Software Engineering, or a related field.
5+ years of experience in software engineering, with at least 3 years focusing on backend systems and ML infrastructure.
Must Have:

Strong past experience with Ray or Kubernetes.
Strong proficiency in Python and at least one systems programming language (Rust, C++ or Go).
Solid understanding of model serving frameworks (e.g., TensorFlow Serving, NVIDIA Triton).
Experience with a ML framework such as TensorFlow, PyTorch, or JAX.
Experience with model compression and optimization techniques.
Strong knowledge of cloud platforms (AWS, GCP, or Azure) and their ML-specific services.
Familiarity with distributed systems and microservices architectures.
Experience with high-performance, low-latency systems.
Ideal candidate will have:
- Experience with GPU programming is a plus.Additional InformationThe role is based in the Bay Area (San Francisco). Candidates are expected to be located near the Bay Area or open to relocation.Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company.
  #J-18808-Ljbffr

Keywords: Genmo Replay, Tracy , Staff AI Infra Engineer (serving API), Engineering , San Francisco, California

Click here to apply!

Didn't find what you're looking for? Search again!

Let San Francisco recruiters find you. Post your resume for free!

Get San Francisco Engineering jobs via email.

View more Tracy Engineering jobs

Other Engineering Jobs

Member of Technical Staff, Growth Engineering Bay Area (United States)
Description: You've mastered growth engineering at a world-class growth org. You're comfortable being scrappy, want to move fast, and care about the customer. You could go get another growth engineering role, or you (more...)
Company: Coframe Inc.
Location: San Francisco
Posted on: 11/3/2024

Senior C++ Research Engineer: Databases & AI Startup
Description: Unum is the deep-tech startup reinventing Data-Lakes for extreme scale and AI You can think of it as Snowflake and OpenAI combined. br We are searching for passionate and competitive Senior C Research (more...)
Company: Unum AI
Location: San Francisco
Posted on: 11/3/2024

Senior Backend Engineer
Description: Income inequality is rising and 1 in 4 Americans have 0 saved for retirement. We understand that most people want to take advantage of their employer benefits-like a 401 k program or employee stock (more...)
Company: Lendtable, Inc
Location: San Francisco
Posted on: 11/3/2024

Salary in Tracy, California Area | More details for Tracy, California Jobs |Salary

Sr Refrigeration Field Service Engineer
Description: Sr Refrigeration Field Service Engineer br br Job ID br br 170706 br br Posted br br 11-Jul-2024 br br Service line br br GWS Segment br br Role type br br Full-time br (more...)
Company: CBRE
Location: San Francisco
Posted on: 11/3/2024

Diesel Technician/Mechanic - Roadside Assistance
Description: Diesel Technician/Mechanic - Roadside Assistance br 53 Morrison Ave., Sacramento, CA 95838 br br Position Summary: br This diesel technician/mechanic position at Penske is focused on providing (more...)
Company: Penske Logistics
Location: Lafayette
Posted on: 11/3/2024

Diesel Technician/Mechanic - Roadside Assistance
Description: Diesel Technician/Mechanic - Roadside Assistance br 53 Morrison Ave., Sacramento, CA 95838 br br Position Summary: br This diesel technician/mechanic position at Penske is focused on providing (more...)
Company: Penske Logistics
Location: Pleasant Hill
Posted on: 11/3/2024

Analytics Engineer San Francisco
Description: Airbyte is the open-source standard for EL T . We enable data teams to replicate data from applications, APIs, and databases to data warehouses, lakes, and other destinations. We believe only an open-source (more...)
Company: Tbwa Chiat/Day Inc
Location: San Francisco
Posted on: 11/3/2024

Diesel Technician/Mechanic - Roadside Assistance
Description: Diesel Technician/Mechanic - Roadside Assistance br 53 Morrison Ave., Sacramento, CA 95838 br br Position Summary: br This diesel technician/mechanic position at Penske is focused on providing (more...)
Company: Penske Logistics
Location: Oakley
Posted on: 11/3/2024

AI SaaS Engineer
Description: Company OverviewDocusign brings agreements to life. Over 1.5 million customers and more than a billion people in over 180 countries use Docusign solutions to accelerate the process of doing business and (more...)
Company: DocuSign, Inc.
Location: San Francisco
Posted on: 11/3/2024

Senior / Staff Cloud Infrastructure Engineer San Francisco, CA
Description: Zip is tackling the 50B TAM space to transform the way businesses manage spend. Our co-founders started Zip YC S2020 because they saw the challenges companies had using outdated 20 year old software (more...)
Company: Tbwa Chiat/Day Inc
Location: San Francisco
Posted on: 11/3/2024

Loading more jobs...

Staff AI Infra Engineer (serving API)

Didn't find what you're looking for? Search again!

Other Engineering Jobs

Log In or Create An Account