T
Talent@ Beta
Nebius

QA Engineer

Nebius · Public · Website

Job Description

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

About the Role

We are looking for a technically strong, hands-on QA Engineer to join our hardware team on-site at ODM factories in Taiwan. This is not a checklist job - we're looking for someone who enjoys digging deep into technical issues, investigating root causes, and taking ownership of complex hardware problems.You'll be the key person ensuring the quality of our servers and racks before they ship, but more importantly, you'll play a critical role in debugging failures, analyzing test data, and working closely with RnD, logistics, and factory teams to continuously improve the process and the product.This is a deeply technical role that blends hardware validation, manufacturing QA, and problem-solving - perfect for someone who understands how servers are built and tested, and wants to make sure every unit that leaves the factory is production-grade.

What You'll Own

  • Technical Investigation & Debugging
  • Investigate complex problems (e.g., high GPU failure rate, power-related test failures), gather logs, run diagnostics, and escalate with context to RnD when needed.
  • Drive root cause analysis across factory teams and internal engineering groups.
  • Document findings and help define preventive actions for recurring problems.
  • Act as the first line of technical escalation for hardware issues discovered during factory QA or internal testing.Engineering Support
  • Participate in new platform bring-up sessions together with the visiting RnD teams during on-site trips to ODM labs.
  • Provide technical support, coordination, and hands-on assistance during the bring-up process.
  • Help ensure early-stage hardware behaves as expected, and escalate integration or platform issues to the relevant teams.On-Site Product QA
  • Perform visual inspections of completed products (servers, racks) before packaging.
  • Define and maintain QA checklists and inspection procedures tailored to different product lines.
  • Verify inventory records at the factory against internal system data (part numbers, serials, configurations).
  • Oversee the product packaging process for compliance with defined standards.
  • Supervise pickup operations: ensure outbound trucks meet shipment conditions and schedules.Failure Rate Monitoring & Analytics
  • Collect failure data from vendor-side burn-in and our own test systems.
  • Analyze failure trends and estimate spare part needs for future datacenter deployments.
  • Use dashboards and structured reporting to communicate insights with QA, engineering, and supply chain teams.Feedback Loop & Quality Improvement
  • Gather and process feedback from datacenters on each delivered batch of equipment:
        * Report on packaging issues, impact sensor triggers, shipping anomalies.
        * Assess rack-level build quality: cabling, bracket alignment, labeling.
        * Log systemic hardware issues (design flaws, infant mortality, recurring failures).
  • Forward the feedback to the teams: logistics, ODM partners, hardware RnD, QA.Test Infrastructure & Validation
  • Assist with deployment and maintenance of test infrastructure on-site.
  • Ensure Nebius post-manufacturing hardware validation tests run smoothly (uptime, monitoring, coordination with support team).
  • Coordinate real-time issue escalation and basic triage with factory and internal teams.Local Insight & Communication
  • Communicate relevant local risks and context (e.g., typhoons, holidays, factory-specific constraints) to our global logistics and hardware teams.
  • Maintain productive relationships with factory staff, logistics providers, and internal stakeholders.

Working Conditions & Tools

  • During production peaks, issues may arise that require fast, hands-on debugging and resolution on-site. Flexibility is expected: you may need to stay late to investigate failures in freshly built batches or arrive early to verify and unblock outbound truck shipments. Rapid response and clear communication with engineering and factory teams are critical during these high-pressure periods.
  • Occasional international travel may be expected to Nebius headquarters in Amsterdam or to datacenters in Europe and the US.
  • Daily work tools involve:
      * Managing workflows and escalation via Jira
      * Writing and maintaining technical documentation in Confluence
      * Using Grafana dashboards for monitoring test environments and system health
      * Operating with several internal inventory and test control systems

What You'll Bring

  • Strong technical background in hardware or systems engineering, able to independently investigate and troubleshoot complex issues with server systems.
  • 5+ years of experience in hardware QA, manufacturing supervision, or server validation.
  • A strong background in R&D is a significant plus.
  • Solid understanding of server and rack hardware: components, layout, cabling, power/cooling, diagnostics.
  • Ability to read and interpret technical documentation (e.g., datasheets, system specs, debug manuals).
  • Solid knowledge of electrical engineering fundamentals (e.g., power specs, grounding, signal integrity).

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI 

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. 

If you need accommodations during the application process, please let us know.

About Nebius

Full-stack AI cloud infrastructure platform for model training, tuning, and deployment. Spun out from Yandex, listed on Nasdaq (NBIS).

View company profile

Role Details

Location
Taiwan
Salary (est. USD)
~$88K - $143K (est. USD)

Estimated based on role seniority, stage (Public) & industry benchmarks.

How is this calculated?
Seniority Mid-level
Base range $80K – $130K
Stage adj. Public (+10%)
Adjusted $88K – $143K
Department
Product & Infrastructure
Type
Full-time
Vertical
AI Infrastructure
Posted
11h ago

Career Tools

You'll be redirected to the company's application page

Get roles like this daily

Join our Telegram channels for curated job alerts