Senior / Staff ML Compiler Engineer
Senior / Staff ML Compiler Engineer
Location: San Jose or Irvine, CA | Full-Time
About Our Client
Our client is a fast-growing fabless semiconductor company building next-generation, energy-efficient domain-specific processors designed for edge AI, wireless communications, advanced radar, computer vision, and autonomous systems.
Founded by industry veterans, our client is developing breakthrough System-on-Chip (SoC) architectures leveraging RISC-V and advanced compute technologies to deliver real-time intelligence at the sensor edge.
This is an opportunity to join a highly innovative engineering team working at the intersection of semiconductor architecture, AI acceleration, and compiler technology.
Position Summary
We are seeking a Senior / Staff ML Compiler Engineer to develop and optimize compiler technologies that unlock the performance of next-generation custom silicon.
This is not a traditional application software engineering role.
The ideal candidate brings deep expertise in compiler backend development, code generation, runtime optimization, hardware/software co-design, and performance tuning for compute-intensive architectures.
This individual will work closely with architecture, silicon, systems, and AI teams to ensure software fully enables our client’s processor roadmap.
Key Responsibilities
Compiler Architecture & Development
- Design, develop, and optimize compiler infrastructure for custom processor and accelerator architectures
- Build compiler backends, optimization passes, code generation workflows, and execution pipelines
- Improve instruction scheduling, register allocation, memory access efficiency, and execution performance
- Develop graph-level and operator-level optimization strategies for AI workloads
- Support compiler enablement for emerging compute architectures
Runtime & Performance Optimization
- Develop runtime systems and execution frameworks for optimized inference and compute performance
- Profile workloads and identify bottlenecks impacting latency, throughput, memory bandwidth, and power efficiency
- Debug performance issues across simulation, emulation, and silicon environments
- Build internal benchmarking and performance analysis tools
Hardware / Software Co-Design
- Partner with architecture and hardware teams on next-generation processor development
- Analyze workload behavior to influence architecture decisions
- Help drive tradeoff analysis involving performance, memory efficiency, latency, and power consumption
- Collaborate with silicon teams during bring-up and optimization cycles
AI / ML Workload Enablement
- Optimize execution of machine learning and inference workloads on custom accelerators
- Support deployment of workloads involving:
- computer vision
- edge AI
- radar
- sensing
- autonomous systems
- signal processing
- Collaborate with AI software teams on model optimization and framework integration
Required Qualifications
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related discipline
- 7+ years of relevant experience (Senior level)
- 10+ years of relevant experience (Staff level)
- Strong C/C++ programming expertise
- Deep experience with compiler development and optimization
- Hands-on expertise with one or more of:
- LLVM
- MLIR
- TVM
- XLA
- GCC
- custom compiler frameworks
- Experience in:
- compiler backend development
- code generation
- optimization passes
- runtime systems
- performance profiling
- Strong understanding of:
- computer architecture
- memory systems
- instruction scheduling
- register allocation
- parallel execution
- low-level performance optimization
Preferred Qualifications
- Semiconductor industry experience
- Experience with AI accelerators, NPUs, DSPs, GPUs, or custom compute architectures
- Knowledge of RISC-V architectures
- Experience with:
- graph optimization
- quantization
- inference optimization
- edge AI deployment
- embedded systems
- HW/SW co-design
- Familiarity with autonomous systems, radar, signal processing, or computer vision workloads
