π» Site Reliability Engineer | DevOps Engineer
π Based in Houston, Texas, United States
π§ Harrimerri10@gmail.com | π 713-480-9493
π Open to Remote Roles
Iβm a Site Reliability Engineer and DevOps Engineer with 8+ years of cloud and production engineering experience, specializing in reliable, scalable, and observable infrastructure across AWS, GCP, and Azure.
My experience spans Kubernetes, Terraform, Linux, CI/CD, observability, infrastructure automation, incident response, and production reliability. At T-Mobile, I work across production readiness, SLI/SLO and error-budget ownership, incident response, release support, capacity planning, recovery planning, and reliability improvements.
Iβm comfortable troubleshooting complex production issues across Kubernetes, Linux, networking, cloud infrastructure, TypeScript services, and Ruby on Rails applications, tracing failures across multiple layers until the underlying issue is understood.
Earlier at Salesforce, I built reusable infrastructure and delivery automation that reduced standard environment provisioning from several days to under one hour, shortened release time by approximately 60%, and contributed to performance improvements of more than 20% in page-load and service response times.
Iβm passionate about building resilient production platforms, automating operational work, improving observability, and turning recurring incidents into durable engineering improvements.
AWS: EC2, VPC, IAM, S3, RDS, ECS, EKS, ALB, Auto Scaling, CloudWatch
- Kubernetes cluster operations
- Amazon EKS
- Docker
- Helm
- Readiness and liveness probes
- Rolling deployments
- Scaling and resource management
- Pod troubleshooting
- Image-pull and configuration troubleshooting
- Kubernetes events and production diagnostics
- Terraform
- Reusable Terraform modules
- Infrastructure as Code
- Ansible
- Python automation
- Bash utilities
- Environment standardization
- Automated health checks
- Log collection and validation
- Repeatable infrastructure changes
- GitHub Actions
- Jenkins
- CI/CD pipelines
- Build and deployment automation
- Release validation
- Rollback planning
- Production release support
- Deployment troubleshooting
T-Mobile
February 2024 β Present | Remote, United States
- Partner with development and platform teams before production changes to review architecture, service dependencies, capacity, observability, rollback plans, and operational risk.
- Own SLI, SLO, and error-budget reviews, using availability, latency, error rates, resource pressure, and incident trends to guide release decisions.
- Review production dashboards and signals around releases and maintenance windows, including health probes, saturation, traffic behavior, dependency health, and recovery readiness.
- Stay engaged after deployments to identify regressions and reliability issues early.
- Investigate live production incidents across AWS, GCP, Azure, Kubernetes, Linux, TypeScript, Ruby on Rails, DNS, TLS, and network dependencies.
- Trace metrics, logs, Kubernetes events, application behavior, and host signals to identify failures across infrastructure and application layers.
- Troubleshoot Kubernetes issues including unhealthy readiness/liveness probes, failed pods, image-pull errors, configuration problems, downstream dependencies, and CPU or memory limits.
- Participate in root-cause and post-incident reviews, separating immediate recovery actions from longer-term engineering improvements.
- Convert recurring failure patterns into improvements to alerts, runbooks, configuration, and deployment practices.
- Build Python and Bash utilities for health checks, log collection, maintenance, and routine validation.
- Use Ansible for repeatable server changes that are safer and easier to review than one-off manual updates.
- Review backup results, failover procedures, capacity trends, access changes, patching plans, and maintenance risks.
- Mentor engineers through production troubleshooting, observability, Kubernetes operations, and reliability reviews.
- Document operational reasoning and troubleshooting approaches to improve knowledge sharing across engineering teams.
- Work across AWS, GCP, and Azure production environments using common operational practices for access, monitoring, deployment safety, and recovery.
Salesforce
March 2018 β February 2024 | Remote, United States
- Provisioned AWS environments using Terraform across VPC, EC2, IAM, S3, RDS, load balancing, Auto Scaling, and monitoring.
- Created reusable infrastructure modules that reduced standard environment provisioning from several days to less than one hour.
- Built reusable Terraform and Ansible patterns for development, staging, and production environments.
- Supported cloud networking and access patterns including VPC design, subnets, routing, security controls, load balancing, IAM, DNS, and TLS.
- Troubleshot connectivity and permission issues spanning application and infrastructure boundaries.
- Built GitHub Actions and Jenkins pipelines for build, testing, and deployment.
- Automated repeated release processes, shortening release time by approximately 60%.
- Containerized three web applications using Docker.
- Supported application deployments through Amazon ECS and Kubernetes.
- Worked with developers during releases and production troubleshooting to distinguish application defects from environment, dependency, and platform issues.
- Administered Ubuntu, CentOS, and RHEL systems, including patching, systemd services, storage, logs, SSL/TLS, user access, and baseline hardening.
- Supported business-critical applications through planned maintenance and production incidents.
- Implemented CloudWatch, Prometheus, and Grafana monitoring for host and application health.
- Used observability data to identify performance bottlenecks, contributing to improvements of more than 20% in page-load and service response times.
- Maintained infrastructure code, deployment notes, runbooks, and environment documentation in version control.
- Built a multi-node Kubernetes environment using EKS/kops and Helm.
- Added readiness and liveness checks.
- Implemented rolling deployments and scaling rules.
- Built Prometheus and Grafana dashboards for monitoring.
- Practiced recovery from failed pods, resource pressure, and configuration changes.
- Created reusable Terraform modules for:
- VPC
- Public and private subnets
- Application Load Balancer
- Auto Scaling
- IAM
- EC2
- RDS
- Used variables and outputs to separate environment-specific configuration while avoiding duplicated infrastructure code.
- Designed backup validation checks and recovery procedures.
- Created controlled failure scenarios to expose service dependencies.
- Validated recovery behavior across infrastructure components.
- Turned findings into clearer operational documentation and runbooks.
- Linux Foundation Certified System Administrator (LFCS)
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Security Specialist (CKS)
- AWS Certified Solutions Architect β Associate
Texas State University
Hands-on coursework covering:
- Linux Administration
- AWS
- Kubernetes
- Docker
- Terraform
- CI/CD
- Python
- Bash
- Networking
- Security
- Monitoring
- Automation
University of Debrecen
Completed two years of undergraduate coursework.
No degree awarded.
Production Readiness
β
SLIs / SLOs / Error Budgets
β
Observability & Monitoring
β
Incident Response
β
Root-Cause Analysis
β
Automation & Reliability Engineering
β
Resilience & Recovery
Reliable systems are built through strong engineering practices, clear observability, thoughtful automation, and continuous learning from production failures.
I focus on making infrastructure and operations:
- πΉ Repeatable
- πΉ Observable
- πΉ Automated
- πΉ Resilient
- πΉ Reviewable
- πΉ Easier to operate
πΌ Open to remote Site Reliability Engineering and DevOps opportunities.
π Houston, Texas | π Remote, United States
π 713-480-9493
π GitHub: Codeprojectingfuture
