SDE II · Platform / DevOps / MLOps · Hyderabad, India
I write the Python services, pipelines and Kubernetes clusters behind production AI systems.
At work I own the platform side of Baseer Builder, a no-code computer-vision product. It runs on 50+ GPU nodes and 2,000+ cameras for a Saudi municipality, and I also deploy it into air-gapped government clusters. In my own time I build tools for awkward networks and Kubernetes on laptops.
$ kubectl get engineer zafeer -o yaml
spec:
writes: [python, bash, typescript]
runs: [kubernetes, argocd, jenkins, vault, ansible, terraform]
ships: [training jobs, notebooks, device fleets, camera pipelines]
status:
phase: Running # since 2024-05, intern → SDE II| Project | What it does |
|---|---|
| causeway | View and record cameras behind VPNs, jump hosts and SSH tunnels. Each customer gets its own network namespace, live view uses WebRTC with HLS fallback, and recordings survive a dropped link. 298 tests. |
| kuiqctl | A single-node kubeadm cluster that stays Ready when your laptop changes Wi-Fi or gets a new DHCP lease. |
| cloudDoc | Private AWS EKS document pipeline: API Gateway → VPC Link → internal ALB, with workers on SQS, files in S3 and Pod Identity instead of static keys. |
| scribblr | A full-stack blog with comments, live notifications and Google/OTP auth on Cloudflare Workers + R2. Live at codesphere.live. |
| k8s-ops-toolkit · k8s-monitoring | Ansible playbooks for k8s operations, and a Prometheus + Grafana monitoring stack. |
| Floating-IP · wireguard-vpn-setup · ssh-tunnel | Small networking tools: Keepalived VIP failover, a WireGuard site link, and RTSP tunnels through jump hosts. |
- Model training. Training runs as a Kubernetes Job with pause, resume and priority, live MLflow metrics and Kafka progress events. PyTorch DDP cut a run from 55h to 5h.
- Notebooks and cameras. Jupyter runs as StatefulSets, each on its own subdomain, and idle notebooks shut down automatically. A MediaMTX RTSP → HLS service handles thousands of cameras.
- Device fleet. A deploy-manager service runs Ansible from k8s Jobs to onboard, patch and deboard GPU machines, reporting a 7-phase lifecycle over Kafka.
- Delivery. Jenkins builds images and bumps tags in an infra repo that Argo CD syncs to dev, prod and EPM. Every secret lives in Vault, and a self-hosted Harbor registry serves 8 teams.
- Platform setup. Moved every service into k8s, cutting full setup from 3 days to 20 min. Built HA NGINX on Keepalived and migrated from InfluxDB to ClickHouse.
- Client deployments. Air-gapped RAG and meeting-bot platforms for a Saudi ministry. An approval-gated Alibaba + Azure DevOps setup for Expro that saves SAR 300K/yr. Multi-VPC ACK for Monsha'at.
A few bugs I've enjoyed killing
- Meeting bots stuck on one node. A shared ReadWriteOnce volume plus a required podAffinity pinned every bot pod to one node. Bots now record to local disk and stream recordings back as a tar over
kubectl exec, so they spread across all 4 workers. - A manifest that would have deleted 22 live secrets. I caught it before apply, shipped a strategic-merge patch instead and reconciled the repo to zero drift.
- The office subnet was blackholed. A Docker bridge was allocating a /16 that overlapped it.
- Deploys silently lost their config. A Vault token expired with nothing renewing it.
Open to Platform, MLOps and AI-infrastructure roles: remote, UAE or KSA. The fastest way to reach me is email.



