Copart, Inc. a technology leader and the premier online vehicle auction platform globally, with over 200 facilities located across the world, Copart links vehicle sellers to more than 750,000 buyers in over 190 countries. We believe in providing an unmatched experience, every day and everywhere, driven by our people, processes, and technology.
Copart, a global leader in online vehicle auctions, is seeking a highly skilled and proactive Site Reliability Engineer Intern to join our SRE team at our Dallas location. This role is pivotal in ensuring the stability, performance, and reliability of Copart’s critical applications and global data center infrastructure through advanced monitoring, troubleshooting, and automation.
As a key member of our 24/7 operations, you will contribute to meeting our stringent SLA commitments and play a crucial role in maintaining operational excellence.
Wha…
What you'd do
- Proactively monitor Copart’s global data centers and application infrastructure using internal tools, identifying and resolving issues before they impact operations.
- Design, build, and optimize monitoring and automation tools using Python and Ansible to collect key metrics and enable automated remediation of infrastructure and application issues.
- Maintain and optimize a suite of monitoring and collaboration tools, including Datadog and Kubernetes-related observability tooling.
- Conduct monthly security patching for operating systems and critical applications.Collaboration & Improvement
- Support for Infrastructure teams, quickly and efficiently troubleshooting and communicating issues across various Copart domains.
- Partner with cross-functional internal teams, including Product Development, DevOps, Network, Systems, and Database teams.
- Develop analytical and reporting capabilities to monitor performance, identify areas for improvement, and implement quality control plans using Data Dog.
- Create and maintain comprehensive standard operating procedures (SOPs), system diagrams, and training materials for team use.
- Support incident management activities including triage, tracking, root cause analysis, and monthly review of lessons learned.
- Process change paperwork throughout the day as part of daily operational requirements.Incident Management Requirements
- Starts
- 2026-08-19