Back to Search Results
Get alerts for jobs like this Get jobs like this tweeted to you
Company: AMD
Location: Singapore, Singapore
Career Level: Mid-Senior Level
Industries: Technology, Software, IT, Electronics

Description



ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we're looking for talent who feel the same: people who want to leave the planet better than they found it, those who don't shy away from humanity's challenges but are determined to help solve them.

 

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you're designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.



THE ROLE:

Returns Debug and RMA execution in Quality & Reliability organization, provide supportive functions to the organization to ensure customer quality issues are being addressed, evoking the required actions via failure analysis to improve product quality.

 

THE PERSON:

You will also need to possess strong verbal and written communication skills, which are essential when working with a global team. A proactive, outstanding teammate who focuses on teamwork, team building, and growing team success.

AMD's environment is fast-paced, results-oriented, and built upon a legion of forward-thinking people with a passion for winning technology!

 

KEY RESPONSIBILITIES:

  • Perform System-level and board-level Failure Analysis (FA) for Data center CPU/GPU products, covering customer returns and field failures.
  • Conduct System-Level Test (SLT) using internal test board or on customer platform to duplicate customer reported failures and isolate the cause of failure.
  • Perform board power up test and functional test to isolate board or component level failure.
  • Collaborate with ASIC Engineering teams for in-depth GPU ASIC/die-level investigations, fault isolation, and root cause analysis.
  • Investigate Excursion and Critical Issues, supporting DPPM (VF/NFF) improvement initiatives.
  • Drive test coverage analysis and enhancements to improve failure detection and mitigation.
  • Partner with validation, firmware, and hardware teams to resolve hardware, software, and platform issues.
  • Innovate, prototype, and evaluate new FA tools to improve GPU failure analysis capabilities.
  • Knowledgeable in functional test and stress software to enhance debug efficiency.
  • Provide technical assistance, resources, and equipment to support engineering teams in testing and debugging activities.
  • Plan, set up, and install server racks with air and liquid cooling capabilities for advanced test infrastructure.
  • Work closely with program managers and product line quality (PLQ)/customer Interacting teams to align failure analysis report writing with external customer communication.
  • Document debug findings, root cause analysis, and corrective actions in clear, concise technical reports.
  • Serve as the local product owner, responsible for tracking and releasing platform screening programs and BKC revision related to server rack level, board level and OSV programs.
  • Act as the Go-To technical expert for owned products, supporting test program contents, FA methodologies, and customer queries.
  • Proactively identify opportunities for process improvement, code quality enhancements, and hardware coverage expansion.
  • Other duties as assigned by supervisor.

 

PREFERRED EXPERIENCE:

  • Strong in either silicon or board level debug or Failure Analysis knowledge.
  • Candidate should be analytical and detail-oriented, strongly interested in debugging complex systems, self-starter, and a fast learner
  • Excellent skill in code development, familiarity with Linux and modern software tools/benchmarks and techniques for development.
  • Understanding of GPU or x86 architecture knowledge is much preferred
  • Knowledge or experience in server or data center hardware or platform is a plus.
  • Experience working with power management features such as POST, P-states, etc.
  • Knowledge of industry standards like PCIE, USB, or high bandwidth memory is a strong plus
  • JTAG knowledge is a plus.
  • Strong understanding of BIOS or memory firmware is a plus.
  • Experience programming experience with C++, C#, Python, HTML, or JAVA.
  • Experience with PC HW debugging, including voltage, networking, storage, and thermal control.
  • Experience with building computer systems (desktops, laptops, servers, etc).
  • Experience in server installation, configuration, and maintenance is a plus.
  • Good analytical and problem-solving skills.
  • Strong understanding of hardware debugging, test equipment (multimeters, oscilloscopes, thermal cameras, etc.), and system-level troubleshooting.
  • Proficient in reading electronic schematics and electronic component datasheets.
  • Proficient in AI tool or machine learning knowledge is a plus. 

 

ACADEMIC CREDENTIALS:

  • Bachelor's or Master's degree in Electrical and Electronic Engineering, Computer Engineering with preferred > 5 relevant years of experience.

 

LOCATION:

Singapore

 

#LI-CK1



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD's “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.


 Apply on company website