Role: Site Reliability Engineer
Experience: Min 10 Years
Locations: Jersey City, NJ & Dallas, TX
Work mode: Hybrid
Employment: W2
The Application Support Engineering role advances Site Reliability Engineering (SRE) practices for applications running in production. The role scope includes Application Support Engineers, SDETs, Software Engineers, and SREs focused on improving reliability, observability, recovery, automation, and operational readiness.
This role brings production support expertise and engineering discipline earlier in the lifecycle to influence design, validate reliability requirements, reduce production risk, and drive measurable operational improvement.
Your Primary Responsibilities
- Participate in design reviews, sprint zero, and delivery planning to define and validate reliability requirements, including resiliency, observability, fault tolerance, performance, scalability, holiday and special-day processing, and disaster recovery.
- Collaborate with Major Release Management to ensure each release meets SRE standards for observability, resiliency, and reliability requirements, support readiness, and knowledge base coverage.
- Define and improve monitoring, observability, dashboards, telemetry coverage, and alert strategy to strengthen outage detection, reduce noise, improve signal quality, and accelerate incident response.
- Assist in major incident response and root cause analysis by identifying observability gaps, improving telemetry and knowledge articles, and driving actions that reduce repeat incidents.
- Drive automation, intelligent tooling, and AI-assisted remediation to reduce manual toil, improve consistency, accelerate recovery, and scale operational support.
- Serve as the operational readiness authority before production releases by validating reliability requirements, assessing support readiness, surfacing production risks, and confirming release supportability.
- Lead capacity, performance, workload trend, and resiliency analysis to ensure applications scale reliably under normal, peak, and stress conditions.
- Establish and track reliability metrics such as availability, incident volume, MTTx, alert quality, automation coverage, reliability requirement compliance, change failure rate, and repeat incident reduction.
- Participate in application reliability governance and service reviews by presenting incident trends, compliance metrics, operational risks, improvement actions, and readiness gaps.
- Prepare executive reporting on reliability posture, release readiness, observability maturity, alert quality, incident trends, automation progress, risks, and improvement outcomes.
- Promote SRE practices through mentoring, standards adoption, best-practice sharing, and approved AI tools that improve knowledge, observability, performance, security, and maintainability.
Qualifications
- Minimum of 10+ years of related technical experience across application support engineering, software engineering, site reliability engineering, production support, or application operations.
- Bachelor’s degree preferred or equivalent practical experience.
- Experience supporting business-critical applications in production environments.
- SRE, observability, automation, or ITIL certifications are a plus.
Talents Needed for Success
- Proven experience in one or more in-scope roles including Application Support Engineer, SDET, Software Engineer, or SRE, with responsibility for improving reliability practices, application validation, observability coverage, automation frameworks, and reliability standards.
- Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation.
- Programming or scripting proficiency in one or more languages such as Python, Java, Go, PowerShell, or similar for automation, tooling, and operational efficiency.
- Familiarity with distributed applications, middleware, messaging, batch processing, real-time processing, and production application behavior in high-availability environments.
- Experience in financial services, capital markets, regulated environments, or other high-availability operational settings.
- Demonstrated participation in disaster recovery, performance testing, resiliency testing, release readiness, incident response, and root cause analysis.
- Knowledge of AI concepts, data platforms, anomaly detection, incident correlation, and intelligent automation use cases.
- Strong collaboration skills across application support, application development, release management, risk, security, business, and vendor stakeholders.
- Ability to translate production support insights into actionable engineering improvements that reduce risk, improve stability, and enhance customer experience.