Skip to content
View HillaryIkhais's full-sized avatar
💭
Polymath
💭
Polymath

Block or report HillaryIkhais

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
HillaryIkhais/README.md

I care about what happens underneath the API call. The retrieval pipeline that decides what the model sees. The evaluation framework that proves whether it actually works. The engineering that makes inference run on hardware that doesn't have a GPU. The agent architecture that doesn't fall apart when two agents disagree about the state of the world.

Most of my work lives at the intersection of ML systems engineering and applied AI — building things that need to work under real constraints, not just in a demo.


Featured Work

Live Domain Model

Agentic behavioral intelligence platform.

Sylon reads raw customer feedback and builds psychological timelines — tracking how taste evolves, where expectations shift, and what actually drives churn. It treats customers as evolving psychological entities, not static segments.

What I built
  • Temporal phase analysis engine that splits user histories into behavioral phases and extracts drift signals
  • Cross-domain translation engine that cold-starts recommendations by mapping psychological drivers across unrelated domains
  • Multi-agent orchestration with intent routing, parallel extraction swarms, and a conversational strategist
  • Zero-shot ranking evaluated against hold-out data
Evaluation metrics
Metric Score
RMSE (rating prediction) 0.7906
NDCG@10 (ranking) 0.1605
HitRate@10 0.2000
ROUGE-L (generation fidelity) 0.1311

Ablation without temporal phase splitting: RMSE degrades to 1.4491, NDCG@10 drops to 0.0652. Static summaries aren't enough.

Python FastAPI Next.js SQLite ElevenLabs

Live Demo →

Benchmarked Domain Model

Offline financial intelligence for African SMEs.

MOVA converts messy WhatsApp messages, OPay SMS, and Pidgin voice-note transcripts into structured financial records — completely offline on an 8GB laptop. No cloud. No API fees. No internet required after model download.

The hard problem

Correctly understanding who owes whom in informal African commerce.

"Chinedu still dey owe me 85k" is a receivable. "I wan pay Alhaji Bello 250k" is a payable.

Mixed Pidgin/English. Implicit context. Multiple transactions per message.

Benchmark (130-example Nigerian economic test set)
Metric Score
Entity extraction 98%
Debt direction 92% (improved from 35% via prompt engineering)
Status detection 91%
Amount extraction 88%
Full record accuracy 78%
On-device Value
Model Llama 3.2 3B (Q4_K_M)
Throughput 7.32 tok/s on Apple M3
Peak RAM 4 GB
Disk 2.02 GB

Llama llama.cpp Tauri Rust TypeScript


Winner Domain AI

Hands-free AI agent for paramedic patient transport.

Paramedics talk. PulseRelay listens, remembers, and hands it off. It extracts structured clinical data from natural speech in real time — vitals, medications, patient demographics — tracks trends, asks for clarification on incomplete data, and generates a complete handoff summary for the receiving hospital.

The critical design choice: Gemini handles understanding language. Deterministic Python code handles everything else — storing values, validating ranges, calculating trends, tracking confidence. No hallucinated vitals. No invented medications. The AI understands; the code decides.

Gemini Google ADK FastAPI Cloud Run Firestore

Competition Domain Runtime

Control layer for autonomous AI agents.

No consequential action should execute simply because the agent is confident. RECKON enforces that principle through an 11-phase state machine with recovery contracts, red-team subagents, and human approval gates.

INTAKE → INVESTIGATION → ANALYSIS → ACTION_PLAN → RECOVERY_CONTRACT
→ SANDBOX_VALIDATION → RED_TEAM → DECISION → HUMAN_CHECKPOINT
→ EXECUTION → VERIFICATION
Action Type RECKON Behavior
Read-only Execute autonomously
Reversible Requires human approval
Destructive Blocked entirely
Unknown Blocked entirely

TypeScript TrueForge MCP Ollama Node.js


Tests Domain Integration

Semantic data-integrity detection for self-healing scrapers.

Self-healing scrapers fix broken extraction logic automatically. But how do you know the repaired scraper is still returning the right data? Vanished is the verification layer — it compares scraper output across snapshots and detects information loss, contract changes, contradictions, and pipeline drift.

A scraper that runs isn't necessarily a scraper you can trust.

Python pytest Bright Data Next.js React

Live Demo →

Live Domain Architecture

Offline-first Edge AI for pharmaceutical verification.

Counterfeit drugs kill over 100,000 people annually in sub-Saharan Africa. CREDO eliminates verification friction by running drug authentication entirely on-device — no internet, no cloud API, no latency.

Built for the hardware that exists in the places where this problem is most acute.

Python Edge AI Offline-first

Live Demo →


How I Think About Building

These aren't principles I wrote on a whiteboard. They're patterns I keep running into.

Evaluation before demos

A system that runs isn't a system that works. I'd rather show you an RMSE score than a screenshot. Every project above has real numbers attached — and where the numbers are weak, I say so.

Abstractions earn their place

I use frameworks when they solve the problem. I drop down to llama.cpp when they don't. MOVA runs on-device because the constraint demanded it, not because offline-first sounds good on a landing page.

Complexity should be provable

Vanished's semantic comparison engine, Sylon's ablation studies, MOVA's benchmark suite — I keep building verification into the system, not after it.

The interesting work is in the plumbing

Retrieval quality, ranking signals, data pipelines, inference optimization, agent coordination — the model is the easy part. Everything around it is the actual engineering problem.


Currently Exploring

Actively Building

Retrieval and ranking systems with real evaluation

Agent coordination that survives multi-turn state

Local inference optimization for constrained hardware

Learning

Knowledge graph memory for agent coherence

Structured evaluation for generative systems

Robotics and embodied AI

Interested In

When agents reason about their own reliability

ML systems that degrade gracefully

Retrieval augmentation meets agent memory


Tech I Work With

Python TypeScript Rust PyTorch
FastAPI Next.js React Docker
Llama llama.cpp PostgreSQL Vercel

Currently building. Open to interesting problems.

GitHub

Pinned Loading

  1. Sylon Sylon Public

    An LLM that reads customer review histories and builds psychological timelines on how their taste evolved, their random contradictions, the shift in what they expect over time.

    Python 1 1

  2. ClarityCall ClarityCall Public

    A privacy-first shield against AI voice scams. ClarityCall runs a genuine neural tensor pipeline natively in the browser to detect and block deepfakes.

    JavaScript

  3. Cordane Cordane Public

    The autonomous consensus engine that surfaces enterprise contract risks before you sign. Powered by a Multi-Agent System.

    TypeScript

  4. CREDO CREDO Public

    An offline-first Edge AI application that eliminates friction in pharmaceutical verification to combat the counterfeit drug crisis.

    Python

  5. FoodGuard FoodGuard Public

    Real-time food safety engine that tracks recalls across supply chains and blocks contaminated products at cash registers.

    TypeScript

  6. Korda Korda Public

    TypeScript 1