Jobiglo

No results.

Senior Site Reliability Engineer (MAAS)

Pragmatike

New Remote
Remote Senior 🇬🇧 English
MAAS Kubernetes Ansible Bash Python Terraform OpenTofu Git Prometheus Grafana Alertmanager VictoriaMetrics VictoriaLogs Proxmox KVM/libvirt OpenStack VMware IPMI Redfish BMC RAID GPU infrastructure VLAN L2/L3 routing Bonding VPN Firewalls DNS

Job description

About the role

Pragmatike is hiring a Senior SRE / Infrastructure Engineer to operate and scale a distributed infrastructure platform spanning bare‑metal GPU nodes, Kubernetes, virtualization, networking, and multi‑site environments. The role is fully remote within EU timezones and involves hands‑on ownership of infrastructure end‑to‑end.

Key responsibilities

  • Operate and maintain large‑scale Linux infrastructure on Debian/Ubuntu bare‑metal and virtualized environments.
  • Own MAAS‑based bare‑metal provisioning, including PXE, commissioning, cloud‑init and API automation.
  • Run production Kubernetes clusters: upgrades, node pools, networking, storage, security hardening and troubleshooting.
  • Design and maintain multi‑site networking (VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, DNS).
  • Automate provisioning and operations with Ansible, Bash/Python, Terraform/OpenTofu and Git‑based workflows.
  • Manage observability platforms such as Prometheus, Grafana, Alertmanager and VictoriaMetrics/VictoriaLogs.
  • Lead incident response, on‑call coverage, post‑mortem analysis and reliability improvements.
  • Handle virtualization platforms (Proxmox, KVM/libvirt, OpenStack, VMware) and GPU passthrough where required.

Required profile

  • 5+ years of hands‑on SRE, infrastructure or platform engineering experience.
  • Expert‑level Linux administration (Debian/Ubuntu) and production experience with MAAS.
  • Deep experience operating Kubernetes in production, including networking, storage and upgrades.
  • Strong network engineering skills across VLANs, routing, bonding, VPNs, firewalls and DNS.
  • Proven automation expertise with Ansible, Bash/Python and Terraform/OpenTofu.
  • Experience with observability tools (Prometheus, Grafana) and on‑call incident response.

Required skills

  • Linux (Debian/Ubuntu)
  • MAAS
  • Kubernetes
  • Ansible
  • Bash
  • Python
  • Terraform / OpenTofu
  • Git
  • Prometheus
  • Grafana
  • Alertmanager
  • VictoriaMetrics / VictoriaLogs
  • Proxmox
  • KVM / libvirt
  • OpenStack
  • VMware
  • IPMI / Redfish
  • BMC
  • RAID
  • GPU infrastructure
  • VLAN, L2/L3 routing, bonding, VPN, firewalls, DNS

What we offer

  • 100% remote work with flexible EU‑based hours.
  • High‑impact role with significant technical ownership and autonomy.
  • Opportunity to shape the architecture and operational foundations of a growing cloud platform.
  • International, engineering‑driven team focused on automation, reliability and scale.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Pragmatike.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 6 hours ago

Expires 1 month from now

3 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Pragmatike