← Tillbaka till jobblistanAnsök / Läs mer →
Software Engineer
ÖppenAurora Engineering AB
incidenthanteringinformationssäkerhetit-teknikdatornätverk
Anställningsform
Tillsvidareanställning (inkl. eventuell provanställning)
Nivå
–
Deadline
1 okt. 2026
Relevanspoäng
30/100
Källa
Arbetsförmedlingen JobStream
Arbetsgivartyp
privat
Sammanfattning
This role carries clear accountability and measurable outcomes in the following areas: 1. End-to-end observability (design → implementation → continuous improvement) 2. Systematic cloud cost optimization across AWS & GCP (FinOps) 3. Production reliability governance and risk reduction 4. Root cause analysis (RCA) and systemic improvement of major incidents You will be expected not only to design but also to deliver, operate, and be assessed against concrete results. 1) End-to-End Observabili
Annons
This role carries clear accountability and measurable outcomes in the following areas: 1. End-to-end observability (design → implementation → continuous improvement) 2. Systematic cloud cost optimization across AWS & GCP (FinOps) 3. Production reliability governance and risk reduction 4. Root cause analysis (RCA) and systemic improvement of major incidents You will be expected not only to design but also to deliver, operate, and be assessed against concrete results. 1) End-to-End Observability What you will own: Independently design and implement a comprehensive end-to-end observability system covering: • Infrastructure (AWS/GCP, Kubernetes, network, storage) • Platform (message queues, databases, caches, API gateways) • Application layer (microservices, critical business flows) • Business layer (key business metrics) You will be expected to produce: 1.Unified Observability Architecture Document • Overall architecture diagram (Metrics + Logs + Traces) • Data flow diagram (collection → processing → storage → visualization) • Tooling selection and justification (e.g., Prometheus, Datadog, OpenTelemetry) 2.Standardized Observability Data Model • Unified metrics naming conventions • Standardized tracing model (Trace ID, Span, sampling strategy) • Structured logging standard (JSON schema) 3.Operational Dashboards • Infrastructure health dashboard • Platform services health dashboard • Business API check of KPI dashboard 4.Alerting System • Defined P0/P1/P2 alert levels • Alert noise reduction strategy • Automated alert routing by team/service 5.SLI / SLO / SLA Framework • At least 5 critical business SLOs defined and tracked • Clear error budget policy 2) Cloud Cost Optimization – FinOps (Core Requirement) What you will own: Lead systematic cost optimization across AWS and GCP without compromising performance, reliability, or user experience. You will implement: 1.Unified Cost Visibility System • Combined AWS + GCP cost dashboards • Cost breakdown by:Team/Product/Service/Environment (Dev/Test/Stage/Prod) 2.Actionable Cost Optimization Plan • Compute (EKS/GKE, EC2/Compute Engine, Serverless) • Storage (S3/GCS tiering, lifecycle policies) • Databases (RDS/Cloud SQL sizing, connection pooling, caching) • Network costs (egress, cross-region traffic) 3.Cost Shift-Left Mechanisms • Cost checks integrated into CI/CD • Mandatory resource ownership and budget limits • Quarterly cost reviews 3) Production Reliability & Incident Governance What you will own: Move from reactive “firefighting” to systematic reliability engineering. Required Deliverables: 1.Incident Management Framework • Standard P0/P1 incident response process • RCA template and follow-up tracking mechanism 2.Reliability Governance Framework • Error budget policy • Standardized canary/gradual rollout process • Automated rollback mechanisms 3.Risk Register • Identified systemic risks and technical debt • Prioritized remediation roadmap 4) Kubernetes & Multi-Cloud Platform Optimization …
Först sedd: 21 sep. 2026 01:02 · Senast sedd: 21 sep. 2026 01:02