Senior / Staff ML Compiler Engineer

San Jose, California
Job TypeDirect Hire
Remote TypeOn-Site

Senior / Staff ML Compiler Engineer

Location: San Jose or Irvine, CA | Full-Time

About Our Client

Our client is a fast-growing fabless semiconductor company building next-generation, energy-efficient domain-specific processors designed for edge AI, wireless communications, advanced radar, computer vision, and autonomous systems.

Founded by industry veterans, our client is developing breakthrough System-on-Chip (SoC) architectures leveraging RISC-V and advanced compute technologies to deliver real-time intelligence at the sensor edge.

This is an opportunity to join a highly innovative engineering team working at the intersection of semiconductor architecture, AI acceleration, and compiler technology.


Position Summary

We are seeking a Senior / Staff ML Compiler Engineer to develop and optimize compiler technologies that unlock the performance of next-generation custom silicon.

This is not a traditional application software engineering role.

The ideal candidate brings deep expertise in compiler backend development, code generation, runtime optimization, hardware/software co-design, and performance tuning for compute-intensive architectures.

This individual will work closely with architecture, silicon, systems, and AI teams to ensure software fully enables our client’s processor roadmap.


Key Responsibilities

Compiler Architecture & Development

  • Design, develop, and optimize compiler infrastructure for custom processor and accelerator architectures
  • Build compiler backends, optimization passes, code generation workflows, and execution pipelines
  • Improve instruction scheduling, register allocation, memory access efficiency, and execution performance
  • Develop graph-level and operator-level optimization strategies for AI workloads
  • Support compiler enablement for emerging compute architectures

Runtime & Performance Optimization

  • Develop runtime systems and execution frameworks for optimized inference and compute performance
  • Profile workloads and identify bottlenecks impacting latency, throughput, memory bandwidth, and power efficiency
  • Debug performance issues across simulation, emulation, and silicon environments
  • Build internal benchmarking and performance analysis tools

Hardware / Software Co-Design

  • Partner with architecture and hardware teams on next-generation processor development
  • Analyze workload behavior to influence architecture decisions
  • Help drive tradeoff analysis involving performance, memory efficiency, latency, and power consumption
  • Collaborate with silicon teams during bring-up and optimization cycles

AI / ML Workload Enablement

  • Optimize execution of machine learning and inference workloads on custom accelerators
  • Support deployment of workloads involving:
    • computer vision
    • edge AI
    • radar
    • sensing
    • autonomous systems
    • signal processing
  • Collaborate with AI software teams on model optimization and framework integration

Required Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related discipline
  • 7+ years of relevant experience (Senior level)
  • 10+ years of relevant experience (Staff level)
  • Strong C/C++ programming expertise
  • Deep experience with compiler development and optimization
  • Hands-on expertise with one or more of:
    • LLVM
    • MLIR
    • TVM
    • XLA
    • GCC
    • custom compiler frameworks
  • Experience in:
    • compiler backend development
    • code generation
    • optimization passes
    • runtime systems
    • performance profiling
  • Strong understanding of:
    • computer architecture
    • memory systems
    • instruction scheduling
    • register allocation
    • parallel execution
    • low-level performance optimization

Preferred Qualifications

  • Semiconductor industry experience
  • Experience with AI accelerators, NPUs, DSPs, GPUs, or custom compute architectures
  • Knowledge of RISC-V architectures
  • Experience with:
    • graph optimization
    • quantization
    • inference optimization
    • edge AI deployment
    • embedded systems
    • HW/SW co-design
  • Familiarity with autonomous systems, radar, signal processing, or computer vision workloads

 

Drag & Drop Resume

(PNG, JPEG, PDF, DOC, TXT)

Message & data rates may apply to all numbers allowed to receive messages

Message frequency varies. Text STOP to opt-out or HELP for assistance