I care about what happens underneath the API call. The retrieval pipeline that decides what the model sees. The evaluation framework that proves whether it actually works. The engineering that makes inference run on hardware that doesn't have a GPU. The agent architecture that doesn't fall apart when two agents disagree about the state of the world.
Most of my work lives at the intersection of ML systems engineering and applied AI — building things that need to work under real constraints, not just in a demo.
|
Agentic behavioral intelligence platform. Sylon reads raw customer feedback and builds psychological timelines — tracking how taste evolves, where expectations shift, and what actually drives churn. It treats customers as evolving psychological entities, not static segments. What I built
Evaluation metrics
Ablation without temporal phase splitting: RMSE degrades to 1.4491, NDCG@10 drops to 0.0652. Static summaries aren't enough.
|
Offline financial intelligence for African SMEs. MOVA converts messy WhatsApp messages, OPay SMS, and Pidgin voice-note transcripts into structured financial records — completely offline on an 8GB laptop. No cloud. No API fees. No internet required after model download. The hard problemCorrectly understanding who owes whom in informal African commerce.
Mixed Pidgin/English. Implicit context. Multiple transactions per message. Benchmark (130-example Nigerian economic test set)
|
|
Hands-free AI agent for paramedic patient transport. Paramedics talk. PulseRelay listens, remembers, and hands it off. It extracts structured clinical data from natural speech in real time — vitals, medications, patient demographics — tracks trends, asks for clarification on incomplete data, and generates a complete handoff summary for the receiving hospital.
|
Control layer for autonomous AI agents. No consequential action should execute simply because the agent is confident. RECKON enforces that principle through an 11-phase state machine with recovery contracts, red-team subagents, and human approval gates.
|
|
Semantic data-integrity detection for self-healing scrapers. Self-healing scrapers fix broken extraction logic automatically. But how do you know the repaired scraper is still returning the right data? Vanished is the verification layer — it compares scraper output across snapshots and detects information loss, contract changes, contradictions, and pipeline drift.
|
Offline-first Edge AI for pharmaceutical verification. Counterfeit drugs kill over 100,000 people annually in sub-Saharan Africa. CREDO eliminates verification friction by running drug authentication entirely on-device — no internet, no cloud API, no latency.
|
These aren't principles I wrote on a whiteboard. They're patterns I keep running into.
|
A system that runs isn't a system that works. I'd rather show you an RMSE score than a screenshot. Every project above has real numbers attached — and where the numbers are weak, I say so. |
I use frameworks when they solve the problem. I drop down to llama.cpp when they don't. MOVA runs on-device because the constraint demanded it, not because offline-first sounds good on a landing page. |
|
Vanished's semantic comparison engine, Sylon's ablation studies, MOVA's benchmark suite — I keep building verification into the system, not after it. |
Retrieval quality, ranking signals, data pipelines, inference optimization, agent coordination — the model is the easy part. Everything around it is the actual engineering problem. |
|
Actively Building Retrieval and ranking systems with real evaluation Agent coordination that survives multi-turn state Local inference optimization for constrained hardware |
Learning Knowledge graph memory for agent coherence Structured evaluation for generative systems Robotics and embodied AI |
Interested In When agents reason about their own reliability ML systems that degrade gracefully Retrieval augmentation meets agent memory |

