Independent infrastructure engineering practice

RetakeData is an independent infrastructure engineering practice. We help teams build and operate systems they can keep under their own control: bare metal, private cloud, hybrid platforms, observability, automation, and private internal tools.

The work is hands-on. We design, implement, test failure modes, document what changed, and stay involved where useful: maintenance, observability, optimization, and training.

Approach

Infrastructure problems rarely stay in one box. A migration can become a storage problem. A storage problem can become an observability problem. An observability problem can expose missing automation, weak runbooks, or a team that was never trained on the system.

migration storage observability automation · runbooks · training

That is where we fit best: crossing the layers without losing the details.

What We Work On

01Foundation

Bare metal, networking, storage, virtualization, Proxmox/Ceph, vSphere/OpenStack migrations, HA design.

bare metalnetworkingstoragevirtualizationProxmox/CephvSphere/OpenStackHA design
02Platform Ops

Automation, deployment workflows, service discovery, observability, alerting, RUM, SLOs, and monitoring as code.

automationdeployment workflowsservice discoveryobservabilityalertingRUMSLOsmonitoring as code
03Private Systems

Internal apps, SRE tools, private AI, RAG, RBAC-aware assistants, and operational tooling that stays inside your infrastructure.

internal appsSRE toolsprivate AIRAGRBAC assistantsstays in your infra
04Open Source Glue

Practical tools around real operations, from SSH fleet access to Terraform providers and audit UIs.

SSH fleet accessTerraform providersaudit UIsreal ops

How We Work

01

Understand first

We start by understanding the current system, not by replacing it. What runs where, who operates it, what breaks, what nobody wants to touch, and what needs to be measurable.

02

Implement with the team

Then we implement with the team: small enough to stay understandable, solid enough to survive production, documented and monitored enough to maintain.

03

Scale when needed

For larger missions, vetted engineers from our network can join the delivery. One point of contact, same operating style.

Background

10+ years across sysadmin · SRE · platform

retakedata is led by Sabri Mjahed, an infrastructure engineer with 10+ years across sysadmin, SRE, and platform engineering.

From racks and virtualization to observability, automation, private AI, and internal tools, the focus stays the same: build systems that are reliable in production and clear for your team to own.

Proof at Scale

Observability

11 TB/day across Loki, Elasticsearch, and Thanos. Grafana deployments serving 2000 users. Ingestion tuning, recording rules, SLO dashboarding, query optimization.

11 TB/day2000 Grafana users400 Loki pods50 TB Thanos
Automation

3000+ VMs delivered across 4 providers. 6-phase pipeline: Git PR, Terraform, Consul/NetBox, Ansible, HAProxy, Centreon. Team-operated, not solo-built.

3000+ VMs4 providersGit to monitored
Virtualization

100+ Proxmox nodes deployed with PXE automation. Production experience on ZFS, NFS, SAN, NVMe-oF, and Ceph. vSphere migrations, HA cluster design.

100+ ProxmoxPXE automationZFS / Ceph / NVMe-oF
Network & Security

DDoS mitigation on F5, distributed firewalling across 1000+ VMs via Ansible. Cisco/Juniper networking, ExtraHop NDR, vulnerability management.

F5 DDoS1000+ VMsdistributed nftables
Sovereign AI

GPU servers running vLLM with private RAG over 1000+ documents. RBAC-aware SRE assistants operating entirely inside client infrastructure.

vLLM6 GPUs1000+ docs RAGRBAC
Databases & Messaging

PostgreSQL HA (Patroni/etcd), large-scale Couchbase, Kafka pipelines, Elasticsearch clusters. Operated at production scale with proper failover and recovery.

PostgreSQL HAPatroni/etcdCouchbaseKafka

Want to scope something?

Bring a platform, migration, observability, or private systems problem. We will help turn it into a concrete plan.