FEL260708P2C2: Senior Platform & DevOps Engineer (Kubernetes | GitOps) - 60% Remote, Brussels, ASAP
joyIT Recruiting GmbH
Contact person: joyIT Recruiting GmbH
Brüssel, Belgium
Energy IndustryObservabilityArgo CDTempoGitOpsKongOpenTelemetryHelmPlatform EngineeringLokiKubernetes SecurityChange ManagementLinuxDevOpsDomain Name System (DNS)Incident ManagementPostgreSQLMongoDBRedisRelease ManagementPrometheusDistributed CachingKusto Query LanguageTCP/IPVulnerability ManagementYAMLMonitoring ResultsDocker ContainerCloud Platform SystemAutoscalingGrafanaInfrastructure as Code (IaC)GitKubernetesRancherHashicorpApi GatewayDevSecOpsDocker
Description
Our client is a leading European operator of critical energy infrastructure. You will join the Platform Operations team and help operate and continuously improve a highly available Kubernetes-based cloud platform supporting mission-critical customer services.
Service Description
The Platform Engineer is responsible for operating and maintaining a highly available Kubernetes-based cloud platform. The focus is on platform operations, lifecycle management, incident response and operational excellence to ensure secure, reliable and performant services.
Primary Objectives
- Maintain platform availability and reliability (SLOs/SLAs)
- Ensure operational readiness across DEV / TEST / ACC / PROD
- Participate in the 24/7 on-call rotation
- Maintain platform observability and security
- Execute platform changes, upgrades and maintenance
Responsibilities
Kubernetes & Runtime Operations
- Operate Kubernetes primitives and platform add-ons (Ingress Controllers, Service Discovery, Workload Identity)
- Troubleshoot Kubernetes failures including pod lifecycle, networking and resource issues
- Execute controlled rollouts with rollback plans
Reliability & Incident Response
- Participate in the 24/7 on-call rotation
- Perform incident triage, mitigation and Root Cause Analysis (RCA)
- Maintain runbooks and operational procedures
Observability & Monitoring
- Operate the observability platform
- Ensure metrics, logs, traces and actionable alerts
- Support incident analysis through telemetry and correlation
Change & Maintenance
- Plan and execute platform changes
- Follow structured change management
- Communicate with stakeholders
- Ensure all changes are documented and auditable
Security & Compliance
- Operate RBAC, network boundaries and secret management
- Apply security updates and patches
- Support vulnerability remediation and audits
Automation
- Automate operational tasks
- Improve standardization and reduce operational risk
- Apply a Platform-as-Code (GitOps) approach
Requirements
Kubernetes
- Multi-cluster lifecycle management
- RBAC & least-privilege design
- Network Policies
- Stateful workloads (CSI, PV/PVC)
- Autoscaling (HPA/VPA)
- Pod Security Standards
- Admission Controllers
- Performance troubleshooting
- Cluster debugging (Networking, DNS, Scheduling, OOM, Crash Loops)
GitOps & Continuous Delivery
- Advanced ArgoCD
- App-of-Apps
- Sync Waves & Hooks
- Drift Detection
- Multi-environment promotion
- Git-based deployments
- Declarative platform design
- Harness.io CI/CD
- Secure secret handling with HashiCorp Vault
Packaging & Configuration
- Advanced Helm
- OCI registries
- Kustomize overlays
Container & Artifact Management
- Docker
- Harbor
- JFrog Artifactory
- Artifact versioning
Security
- HashiCorp Vault
- CSI Integration
- Image vulnerability scanning
- TLS & certificate lifecycle
- RBAC governance
Observability
- OpenTelemetry
- Prometheus or VictoriaMetrics
- Loki
- Tempo
- Grafana
- SLI/SLO concepts
- Error budgets
- Alert optimization
Networking
- TCP/IP
- DNS
- TLS / mTLS
- Kubernetes Services
- Ingress
- Reverse Proxy
- API routing
- Network Policies
Automation
- Advanced Bash scripting
- Infrastructure automation mindset
Nice to Have
- Kong API Gateway
- Redis
- PostgreSQL
- MongoDB
- Kargo (ArgoCD)
Operational Skills
- Production platform operations experience
- Strong troubleshooting across distributed systems
- Calm under pressure
- Excellent communication
- Ability to balance operations with continuous improvement
Ways of Working
- Structured, detail-oriented and risk-aware
- Strong ownership mentality
- Collaborative with Development, Security and Product teams
- Documentation-first mindset
Project Details
Start: August
Duration: until December 2027, 17 months, with possible extension
Location: Brussels, Belgium
Work Model: Hybrid (2 days onsite per week)
Workload: Full-time
Please activate the "Send E-Mail" option when submitting your application. After your application has been submitted, you will automatically receive a link to complete the next steps in the recruiting process.
The further application process for projects is managed via our website: www.joyit.de
We look forward to receiving your application.
Your joyIT Team