54 دقیقه پیش | کد آگهی: 11803105
دستهبندی شغلی
موقعیت مکانی
تحصیلات
محل فعالیت
-
مزایا
-
مهارت ها و زبان ها
نوع همکاری
سایر اطلاعات
Description
Pegah is the technology group driving a wide range of digital products and businesses including Cafe Bazaar Tapsell Metrix Bazaar Pay Metis AI Gapify Athena AI Bebin TV Beeptunes Footballi and more serving over 50 million active users
The Technology Team builds and operates the shared infrastructure platforms services and core capabilities that enable engineering teams across the group to operate reliably at massive scale every single day
We are looking for a Senior Site Reliability Engineer SRE to join our Technology team and help improve the reliability scalability and operational excellence of our production systems across the group
Position Summary
As a Senior Site Reliability Engineer SRE you will work closely with engineering teams to improve the reliability availability scalability and performance of business critical production services You will help define and improve reliability practices participate in incident response and on call rotations reduce operational toil through automation and proactively identify risks before they impact users
This role is a good fit for someone who enjoys understanding complex production systems troubleshooting under pressure collaborating across engineering domains and continuously improving how services are operated and supported
What You Will Do
Maintain and improve a diverse range of production services including in house and open source systems
Work closely with product and engineering teams to understand service behavior dependencies failure modes and operational risks
Participate in on call rotations incident response troubleshooting and production support for critical systems
Troubleshoot complex production issues across applications and their underlying dependencies
Monitor system behavior and proactively identify reliability capacity availability and performance risks
Define measure and improve Service Level Indicators SLIs Service Level Objectives SLOs and other reliability targets
Lead or contribute to incident analysis postmortems and corrective actions to prevent recurring failures
Reduce operational toil through automation better workflows and elimination of repetitive manual work
Improve service observability through meaningful metrics logs traces dashboards and alerts
Review system architecture and production readiness bringing reliability considerations into service design and evolution
Improve resilience through better failure handling graceful degradation recovery mechanisms and capacity planning
Turn incident learnings into better runbooks reliability practices tooling and operational standards
Collaborate across product platform infrastructure delivery and security to continuously improve production reliability and operational excellence
What We Expect
4 years of experience in Site Reliability Engineering Production Engineering or a similar production focused role
Strong problem solving and troubleshooting skills in complex production environments
Solid knowledge of Linux systems and underlying concepts
Good understanding of networking fundamentals TCP IP DNS load balancing and proxies
Good understanding of distributed systems failure modes high availability fault tolerance and recovery patterns
Solid understanding of core SRE concepts such as SLIs SLOs error budgets toil and service reliability
Proven experience owning production systems with a high degree of autonomy driving incidents through to resolution and collaborating effectively across engineering teams
Hands on experience with Kubernetes and containerized workloads in production
Experience with observability systems such as Prometheus Grafana VictoriaMetrics or similar technologies
Experience with centralized logging systems such as Loki or Elasticsearch ELK along with distributed tracing and instrumentation using OpenTelemetry or similar technologies
Good understanding of service mesh and traffic management concepts including the roles of service meshes such as Istio or Linkerd and proxies such as Envoy
Proficiency in Go or Python plus solid shell scripting skills
Strong mindset for automation and operational efficiency
Willingness to participate in scheduled on call rotations and incident response
Nice to Have
Experience defining and managing SLIs SLOs and error budgets in practice
Hands on experience operating service meshes such as Istio or Linkerd or proxies such as Envoy in production
Familiarity with Infrastructure as Code and GitOps tooling such as Terraform Ansible Helm or Argo CD
Experience with software delivery tooling and workflows including GitLab CI CD runners and artifact repositories such as JFrog Artifactory or Nexus
Experience operating or supporting highly available databases messaging systems stateful services or distributed data platforms such as PostgreSQL MySQL or Kafka
Practical experience using AI assisted engineering tools such as Claude Codex Gemini Cursor etc to accelerate troubleshooting automation documentation and day to day technical workflows
شرکت تپسل Tapsell
دسته بندی IT DevOps Server
نحوه همکاری تمام وقت
نوع همکاری حضوری
مدرک تحصیلی کارشناسی
سابقه کار سه تا شش سال
حقوق توافقی
جنسیت مهم نیست
مهارت ها DevOps عیب یابی Linux Grafana SRE
شهر تهران تهران
جویا کار این آگهی را از سایت
جابینجا
استخراج نموده است و هیچ مسئولیتی در قبال این آگهی ندارد.
دقت نمایید که کارفرما حق دریافت هیچ گونه وجهی از کارجو را نداشته و این امر خلاف قانون است. در صورت مشاهده این موارد یا سایر تخلفات با کلیک روی (گزارش آگهی) ما را در ارائه خدمات بهتر یاری نمایید.
در غیر این صورت میتوانید با کلیک بر روی دکمه "درج نظر" نظر خود را در مورد این آگهی ثبت کنید.
جهت اشتراک در شبکه های اجتماعی روی کلیدهای زیر کلیک کنید
همچنین میتوانید لینک کوتاه زیر را جهت دسترسی به صفحه فوق برای اشتراک گذاری کپی کنید
کپی کردن لینک
نظرات در مورد این آگهی: درج نظر