Platform Operations Engineer

Date:  Aug 14, 2026
Location: 

Houston, TX, US, 77010

Company:  NRG

As an NRG employee, we encourage you to take charge of your career and development journey. We invite you to explore exciting opportunities across our businesses. You’ll find that our dynamic work environment provides variety and challenge. Your growth is key to our ongoing success—take the lead in shaping your career development, goals and future!


 

Position Summary 

 

 

The Platform Operations Engineer is responsible for coordinating and improving the operational effectiveness of the Home Services technology platform. This role partners across Engineering, Product, Architecture, Quality Assurance, and Platform teams to improve platform reliability, operational visibility, and service performance through coordination, reporting, operational excellence, and continuous improvement. 

The Platform Operations Engineer supports production operations, operational reporting, environment coordination, monitoring and observability, and AI-enabled operational capabilities to ensure engineering teams have the visibility, processes, and operational support necessary to deliver reliable technology solutions. 

 

 

Key Responsibilities 

 

 

Production Support & Operations 

 

 

  • Coordinate production support activities across Engineering, Product, and Platform teams to facilitate timely incident resolution and effective communication. 

  • Support incident management processes by coordinating investigations, root cause analysis activities, corrective actions, and post-incident follow-up with the appropriate engineering teams. 

  • Coordinate operational readiness activities for releases, maintenance events, and platform changes. 

  •  Track recurring operational issues and coordinate continuous improvement initiatives with responsible teams. 

  • Operational Metrics & Performance 

  • Develop, maintain, and communicate operational dashboards, KPIs, and service performance metrics. 

  • Analyze operational data to identify trends, risks, and opportunities for operational improvement. 

  • Coordinate recurring operational reviews and provide visibility into platform health, service levels, and operational performance. 

  • Support the definition, measurement, and reporting of operational objectives and service quality indicators. 

 

 

Monitoring & Observability 

 

 

  • Partner with engineering teams to ensure appropriate instrumentation, logging, monitoring, and alerting are implemented across platform services. 

  • Identify gaps in observability and coordinate improvements with engineering teams. 

  • Support adoption of monitoring standards and operational reporting practices that improve proactive issue detection and platform visibility. 

  • Promote consistent telemetry and operational reporting across platform services. 

  • Environment Management & Coordination 

  • Coordinate environment planning, scheduling, availability, and utilization across multiple  

  •  Facilitate environment requests, refreshes, deployments, and conflict resolution activities. 

  • Maintain visibility into environment readiness, dependencies, and operational risks. 

  • Communicate environment status, planned activities, and potential impacts to stakeholders. 

 

 

AI & Operational Innovation 

 

 

  • Identify opportunities to leverage AI and automation to improve production support, operational efficiency, and platform reliability. 

  • Partner with engineering teams to implement AI-assisted operational workflows, monitoring, and support processes. 

  • Support adoption of AI-enabled operational tools that improve issue detection, operational insights, knowledge management, and engineering productivity. 

  • Evaluate emerging AI capabilities and recommend practical applications that enhance platform operations. 

 

 

Operational Excellence 

 

 

 

  • Support development and maintenance of operational documentation, runbooks, standard operating procedures, and knowledge resources. 

  • Identify opportunities to improve operational processes through automation and standardization. 

  • Coordinate operational improvement initiatives that enhance platform reliability and support efficiency. 

  • Support operational governance by providing reporting, metrics, and operational insights. 

 

 

 

Cross-Functional Collaboration 

 

 

  •  Partner with Platform Architects, Solution Engineers, Engineering teams, Quality Assurance, and Delivery teams to improve operational effectiveness. 

  • Coordinate cross-functional activities, dependencies, and communications impacting platform operations. 

  • Provide visibility into operational risks, dependencies, and platform readiness to support informed decision-making. 

  • Foster collaboration and continuous improvement to enhance operational maturity and service performance. 

 

 

Required Skills & Experience 

 

 

 

Minimum Requirements 

 

 

  • Bachelor's degree in Information Systems, Computer Science, Engineering, or a related field, or an equivalent combination of education and relevant work experience. 

  • Five (5) or more years of experience in Platform Engineering, Application Support, IT Operations, DevOps, Cloud Operations, Site Reliability Engineering, or a related technical discipline. 

  •  Experience supporting mission-critical production applications within an enterprise environment. 

  • Experience coordinating production support activities across multiple Engineering and business teams. 

  • Experience developing, monitoring, and reporting operational KPIs, dashboards, and  

  • Experience with application monitoring, logging, observability, and alerting platforms. 

  • Experience coordinating environment management activities, including planning, release readiness, deployments, refreshes, and cross-team scheduling. 

  • Experience supporting cloud platforms and enterprise applications, preferably Microsoft Azure and/or AWS. 

  • Experience working with IT Service Management (ITSM) processes, including Incident, Problem, Change, and Release Management. 

  • Strong analytical, troubleshooting, organizational, and problem-solving skills with the ability to identify trends and recommend operational improvements. 

  • Excellent verbal and written communication skills with the ability to collaborate effectively across technical and business teams. 

  • Experience with enterprise operational tools such as Azure DevOps, GitHub, ServiceNow, Azure Monitor, Application Insights, Datadog, Splunk, or similar platforms. 

  • Experience applying AI, automation, or scripting to improve operational efficiency, monitoring, reporting, or support processes. 

 

NRG Energy is committed to a drug and alcohol-free workplace. To the extent permitted by law and any applicable collective bargaining agreement, employees are subject to periodic random drug testing, and post-accident and reasonable suspicion drug and alcohol testing. EOE AA M/F/Vet/Disability. Level, Title and/or Salary may be adjusted based on the applicant's experience or skills.

Official description on file with Talent.

We support the use of AI tools to help you prepare for your interview (e.g., practicing responses, researching the role, or refining your resume). However, during interviews and assessments, we expect responses to reflect your own thinking, experience, and communication. Use of AI to generate or read answers in real time, complete assessments, or misrepresent your qualifications is not permitted and may impact your candidacy.


Nearest Major Market: Houston