The board
- Hours
- Full time
- Where
- Austin, TX
- Posted
- Aug 31
Our vision is to transform how the world uses information to enrich life for all.
Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever.
The Advanced Systems Research and Engineering team develops innovative hardware and software technologies that enable the next generation of Artificial Intelligence infrastructure. The team collaborates closely with engineering, architecture, product, and research organizations to evaluate emerging AI workloads and drive advancements in memory, storage, interconnects, and distributed computing platforms.
What you'd do
- Develop and enhance systems software, profiling tools, and experimentation frameworks for LLM training, LLM inference, and Agentic AI workloads.
- Design, implement, and evaluate memory- and state-management techniques, including caching, tiering, compression, eviction, and lifecycle management for AI serving environments.
- Characterize and optimize AI workload execution across GPUs, CPUs, memory subsystems, storage, and distributed infrastructure, with a focus on latency, throughput, scalability, and resource utilization.
- Build benchmarking, simulation, and automation capabilities to evaluate data placement, migration, scheduling, and performance behavior across heterogeneous memory systems.
- Collaborate with engineering and research teams to develop representative AI workloads, analyze experimental results, and contribute to technical publications, intellectual property, and future platform designs.Minimum Qualifications
- Currently pursuing a Master's or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- Demonstrated experience with AI systems, machine learning systems, computer systems research, or systems software development through coursework, research, or projects.
- Understanding of Large Language Models (LLMs), including transformer execution, attention mechanisms, KV cache behavior, batching, token-level latency, throughput, and memory performance considerations.
- Proficiency in Python and C/C++, with hands-on experience developing, debugging, and optimizing software in Linux environments.
- Experience using GPU-based performance analysis tools and at least one modern AI framework or serving stack, such as PyTorch, vLLM, TensorRT-LLM, NVIDIA Dynamo, or related technologies.Preferred Qualifications
- Starts
- 2026-08-31