
Responsibilities
About the Team The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across ByteDance’s self-built data centers in the United States and Europe. Our scope covers new hardware introduction, data center delivery, production operations, hardware maintenance, configuration and firmware changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse. The team serves as a central coordination point between multiple functions, including: - Hardware New Product Introduction (NPI) - Server and data center operations - Field maintenance and infrastructure management - Hardware vendors and service providers - Supply chain and asset management - Infrastructure platform and automation engineering teams Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle. Role Overview We are looking for a hands-on Production Systems Engineer with a strong foundation in Linux systems, server infrastructure, automation, and production operations. This role is open to engineers across a range of experience levels, from early-career engineers with strong technical fundamentals to experienced infrastructure engineers who can take ownership of complex systems and large-scale initiatives. The scope and level of ownership will grow with experience, ranging from hands-on infrastructure engineering and automation development to leading complex global infrastructure initiatives across organizational boundaries Responsibilities - Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers. - Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency. - Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues. - GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements. Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making. - Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement. - Technical Documentation: Create and improve technical documentation, standard operating procedures, troubleshooting guides, and internal knowledge bases. - Cross-functional Collaboration: Work with infrastructure engineers, hardware teams, data center operations, platform developers, supply chain teams, and other stakeholders on global infrastructure projects.
Qualifications
Minimum Qualification(s) - Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field. - 2 years of experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands-on project experience. - Strong foundation in Linux system administration and troubleshooting, with an understanding of basic server architecture, operating systems, storage, networking, and hardware management concepts. - Programming or scripting experience in Python, Bash, Go, or another modern programming language, with the ability to develop tools or automation for infrastructure or operational tasks. - Hands-on experience troubleshooting system, hardware, storage, networking, or performance-related issues in Linux-based environments. - Strong analytical and problem-solving skills, with the ability to learn unfamiliar technologies quickly and investigate complex technical issues in a structured manner. - Good communication and collaboration skills, with the ability to work effectively with engineers and cross-functional stakeholders across different technical domains and regions. Preferred Qualification(s) - Familiarity with technologies such as BIOS/UEFI, BMC, firmware, PCIe, NVMe, NICs, or hardware telemetry. - Proficiency in Python, Go, Bash, or another programming language for production-grade infrastructure automation, including experience designing, building, or maintaining tools and platforms used in large-scale production environments. - Deep knowledge of Linux administration and troubleshooting, preferably Debian or Ubuntu, combined with strong understanding of server architecture and management technologies such as kernels, drivers, BIOS/UEFI, BMC, Redfish, firmware, PCIe, NVMe, NICs, DPUs, hardware telemetry, and failure diagnostics. - Proven hands-on experience introducing and productionizing large-scale GPU infrastructure, including ownership of hardware NPI or fleet onboarding across qualification, system integration, deployment, production validation, operational handoff, and post-launch reliability. - Strong understanding of distributed AI workload behavior and performance analysis, including collective communication, multi-node training, inference serving, GPU scheduling, checkpointing, workload-related bottlenecks, DCGM, NCCL testing, CUDA profiling, and network-fabric telemetry. - Experience building and operating monitoring, telemetry, hardware management, or automated remediation platforms at substantial scale, with measurable improvements in fleet availability, deployment efficiency, incident reduction, operational efficiency, or reliability. - Experience working directly with OEMs, ODMs, component suppliers, or GPU platform vendors throughout qualification, technical escalation, root-cause analysis, and corrective-action processes. - Experience with one or more advanced infrastructure technologies or engineering areas, such as containerisation and orchestration (e.g., Docker, Kubernetes), infrastructure automation frameworks (e.g., Ansible), AI-powered automation, AI agents, Large Language Models, Retrieval-Augmented Generation (RAG), open-source infrastructure projects, technical publications, patents, or relevant industry standards.
Job Information
About Us
Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.
Why Join ByteDance
Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.
As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.