The board
- Where
- San Jose
- Posted
- Aug 3
Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.
Job Summary
Join our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware.
What you'd do
- Implement the data collection framework for hardware performance counters on a custom PCIe-based accelerator.
- Develop a user-space service for low-overhead tracing of accelerator activity.
- Design and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.
- Create an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.
What they want
- Strong programming skills in C++ or Rust. Experience with Python is a plus.
- Solid understanding of computer architecture, including CPUs, GPUs or AI accelerators, memory hierarchies, and parallel programming.
- Experience or strong interest in low-level performance analysis, profiling, and performance optimization.
- Familiarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto, or similar tools is a plus.
- Experience or strong interest in operating systems, compilers, firmware, drivers, or other low-level systems software.
- Passion for understanding how complex systems behave under real workloads and building tools that help other engineers optimize performance.
- Strong problem-solving skills and curiosity to learn quickly in a fast-paced engineering environment.
- Direct experience developing performance analysis or debugging tools.
- Experience with ML accelerator architectures (GPUs, TPUs, etc.).
- Experience with kernel-mode driver development (Linux or Windows).