Arivazhagan Pandiyan
Open to SRE Lead Roles

Arivazhagan Pandiyan

Site Reliability Engineering professional with 11+ years of experience supporting large-scale production environments across financial services and logistics platforms.

arivu.p@live.in +1 346-599-0347 Houston, Texas, USA www.arivu.site

About Me

I am a seasoned Site Reliability Engineering Lead / Senior SRE with over 11 years of experience supporting large-scale production environments across financial services and logistics platforms. I specialize in building fault-tolerant, scalable infrastructure and ensuring operational excellence for mission-critical applications.

My expertise spans Microsoft Azure and Google Cloud Platform (GCP), with deep hands-on experience in Kubernetes (GKE) environments and enterprise observability platforms including Datadog (APM, Tracing, DBM, RUM), Splunk, and Dynatrace. I serve as the primary Datadog specialist, managing end-to-end observability pipelines and driving telemetry hygiene through OpenTelemetry adoption.

As an Incident Commander, I lead triage for high-severity outages and drive post-incident root cause analysis (RCA) to ensure continuous improvement. I've successfully reduced alert noise from 40% to 10% by implementing AI-driven alerting, Watchdog anomaly detection, and SLO-aligned monitoring standards based on the “Golden Signals” framework.

I'm proficient in Infrastructure as Code using Terraform and Ansible, with strong scripting skills in Python, PowerShell, and Shell. I hold the Microsoft Certified: Azure Administrator Associate (AZ-104) and Datadog Fundamentalscertifications. I'm passionate about reducing toil through automation, optimizing cloud costs, and mentoring engineering teams on observability best practices.

11+
Years Experience
75%
Alert Noise Reduced
15+
Team Members Led
2
Certifications

Technical Proficiency

Certifications

DD

Datadog Fundamentals Certification

Datadog · Verified

MS

Microsoft Certified: Azure Administrator Associate

Microsoft · AZ-104

Cloud & Kubernetes

Google Cloud Platform (GCP)Microsoft AzureGoogle Kubernetes Engine (GKE)KubernetesConfigSync GitOpsHelmDocker

CI/CD & GitOps

Jenkins on KubernetesJCasCStakater ReloaderGitOps Environment PromotionGitHub Actions

Observability

Datadog (APM, Tracing, DBM, RUM)OpenTelemetryELK StackSplunkDynatraceAzure Monitor

Automation / IaC

TerraformAnsiblePythonPowerShellBash / Shell Scripting

Secrets & Security

HashiCorp VaultExternal Secrets OperatorKubernetes SecretStoresRBAC & SecurityContext

SRE Practice

Error BudgetsChaos Engineering (Chaos Mesh)Blameless Post-mortemsGolden SignalsSLIs / SLOs

ITSM / Incident

PagerDutyOpsgenieServiceNowIncident Command & RCA

Platforms & Infra

Linux (RHEL / Ubuntu)VMwareWindows ServerWebLogic TMS

Professional Experience

Site Reliability Engineer

Aug 2025 – Present

Izen Labs (Client: Uber Freight) · Remote

  • On-Prem to GCP Migration: Spearheaded the comprehensive migration of monitoring infrastructure from legacy on-premises data centers to Google Cloud Platform (GCP), ensuring zero observability gaps.
  • Datadog Observability Transformation: Re-architected the monitoring landscape by migrating legacy ELK stack logs and on-prem metrics into Datadog, centralizing telemetry for high-scale GCP workloads.
  • High Availability & Error Budgets: Engineered system reliability to maintain 99.99% availability for mission-critical logistics platforms; managed Error Budgets to balance feature velocity with production stability.
  • Infrastructure as Code (IaC): Utilized Terraform to architect and provision multi-region Kubernetes (GKE) clusters, implementing modular code to ensure consistent environment states.
  • Chaos Engineering Implementation: Enhanced system resilience by conducting scheduled fault-injection experiments using Chaos Mesh, successfully identifying and mitigating single points of failure.
  • Kubernetes Scalability: Optimized application performance via manual and automated scaling (HPA/VPA) of workloads within Kubernetes to handle unpredictable traffic spikes.
  • CI/CD Pipeline Management: Orchestrated automated deployment pipelines using Jenkins and GitHub, incorporating automated testing to reduce deployment-related incidents.
  • Incident Command & RCA: Acted as Incident Commander for high-severity production outages, coordinating cross-functional teams and utilizing ServiceNow for Root Cause Analysis (RCA).
  • Alert Optimization: Leveraged Datadog Watchdog and AI-driven alerting to reduce alert noise from 40% to 10%, focusing the team on actionable events.

Site Reliability Engineering Lead

Aug 2023 – Jul 2024

New American Funding

  • Strategic Team Leadership: Led a high-performing 15-member SRE team responsible for the 24/7 reliability of enterprise-level financial and mortgage systems.
  • Enterprise Monitoring Overhaul: Designed and implemented a comprehensive enterprise monitoring strategy utilizing Azure Monitor and PagerDuty, increasing visibility into legacy applications.
  • Service Level Definition: Established robust service level indicators (SLIs) and service level objectives (SLOs) to align IT performance with business expectations.
  • Operational Automation: Significantly reduced manual toil by automating repetitive operational tasks and diagnostic workflows using PowerShell and Python scripting.
  • Cross-Team Incident Coordination: Managed high-severity production responses, facilitating blameless post-mortems and coordinating long-term stability fixes.

Technology Operations Associate

Oct 2017 – Oct 2022

Wells Fargo India Solutions

  • Infrastructure Maintenance: Maintained enterprise infrastructure supporting mission-critical banking applications.
  • Health Monitoring: Monitored performance across Windows and VMware environments, automating diagnostics with PowerShell.
  • Collaboration: Partnered with network and database teams to resolve complex production incidents.

System Administrator

Sep 2015 – Oct 2017

NTT DATA

  • Server Management: Managed Windows and Linux production servers in a 35-member command center for financial clients.
  • Hardware Operations: Resolved hardware issues via iLO, DRAC, and SMH; managed security compliance and tool integration.

Support Engineer

Nov 2014 – Sep 2015

Firstsource

  • Production Support: Monitored server stability (CPU/Disk/Memory) and managed VSS backups for United Health Care and GHX.
  • Technical Support: Resolved infrastructure alerts and handled incident ticketing through Kayako and XSmart-control.

Assistant Engineer

Jun 2013 – Aug 2014

Cliptos Technologies

  • Systems Deployment: Installed and upgraded healthcare IT systems (Meditos) and managed asset inventory.
  • Field Support: Configured engineering software solutions and assisted in resolving escalated technical issues.

Key Projects & Initiatives

Observability & AIOps

On-Prem to GCP Telemetry & Observability Migration

Izen Labs · Client: Uber Freight

Spearheaded comprehensive migration of legacy on-premises logging and monitoring infrastructure to Google Cloud Platform with centralized Datadog telemetry.

  • Migrated legacy ELK stack logs and on-prem metrics into Datadog, ensuring zero observability blind spots for high-scale GCP workloads.
  • Established OpenTelemetry standard pipelines for distributed tracing across cloud microservices.
  • Provisioned repeatable monitoring infrastructure using Terraform IaC modules.
GCPDatadog (APM/Tracing)OpenTelemetryTerraformELK Stack
Cloud & K8s

Multi-Region Kubernetes & Chaos Resilience Engineering

Uber Freight Workloads

Architected fault-tolerant multi-region Kubernetes (GKE) clusters with automated scaling and proactive chaos engineering fault-injection tests.

  • Provisioned multi-region GKE clusters using modular Terraform code for immutable infrastructure.
  • Conducted scheduled fault-injection experiments using Chaos Mesh to uncover and mitigate single points of failure.
  • Implemented automated HPA/VPA scaling policies to smoothly absorb unexpected freight traffic surges while maintaining 99.99% availability.
Kubernetes (GKE)Chaos MeshTerraformHPA / VPADocker
Observability & AIOps

AI-Driven Alert Optimization & AIOps Platform

Production Operations

Engineered an AIOps telemetry framework that reduced on-call alert fatigue by 75% and automated incident triage workflows.

  • Leveraged Datadog Watchdog AI anomaly detection and Golden Signals framework to drop alert noise from 40% down to 10%.
  • Defined actionable SLIs/SLOs and managed Error Budgets to balance feature release velocity with system stability.
  • Integrated automated incident dispatch and root cause analysis (RCA) reporting through ServiceNow and PagerDuty.
AIOpsDatadog WatchdogPagerDutyServiceNowSLIs / SLOsPython
Observability & AIOps

Enterprise Azure & Dynatrace Observability Overhaul

New American Funding

Designed and executed end-to-end observability strategy across critical financial and mortgage platforms.

  • Configured Dynatrace monitors, Real User Monitoring (RUM), and deep distributed tracing for legacy and modern financial apps.
  • Automated diagnostic and remediation runbooks using PowerShell and Python scripts to eliminate repetitive manual toil.
  • Led 15-member SRE team through high-severity incident responses and blameless post-mortem reviews.
Microsoft AzureDynatraceAzure MonitorPowerShellIncident Command
Automation & AI

AI Tooling & Workflow Automation via MCP

SRE Automation Initiative

Built custom Model Context Protocol (MCP) integrations and AI-assisted workflows to accelerate SRE diagnostics and incident resolution.

  • Connected developer agents (Claude, Cursor) to internal monitoring telemetry and infrastructure APIs via MCP standards.
  • Automated log parsing, anomaly summarization, and routine operational runbooks, reducing MTTR for common incidents.
  • Authored custom Python automation utilities for cloud health audits and configuration consistency checks.
Model Context Protocol (MCP)Claude / CursorPythonAIOpsShell
Automation & AI

Operations & Reliability Analytics Platform

Enterprise Infrastructure & Logistics

Built interactive Power BI and SQL-driven observability analytics dashboards tracking mission-critical operational KPIs and SLA adherence.

  • Constructed executive Power BI dashboards visualizing multi-service SLA adherence, MTTA/MTTR metrics, and incident trends.
  • Aggregated database and operational telemetry using Datadog DBM, Azure SQL, and ServiceNow APIs.
  • Delivered real-time operational insights for both engineering leadership and business operations stakeholders.
Power BIPower QueryAzure SQLDatadog DBMSLA Analytics

Education

Bachelor of Technology (Information Technology)

Sri Venkateshwara College of Engineering, Anna University — Chennai, India

Get in Touch

I'm currently open to new opportunities as a Site Reliability Engineer Lead. Whether you have a question or just want to connect, feel free to reach out!

Location

Houston, Texas, USA

Connect on LinkedIn
Resume

Arivu's Assistant

Hello! I'm Arivu's AI assistant. How can I help you today?