LLM을 실제 업무에 붙이고, 틀린 출력이 실행되기 전에 잡아냅니다.
Backend engineer at 한국리서치 (2024.09 – present). I ship LLM and agent systems into real workflows, and the part I care about is the part that comes after generation — the gate, the trace, the audit line that tells you whether the answer can be trusted.
Search-Pro — Text-to-SQL in daily use at work
C# · ASP.NET Core · Python re-implementation
Every data question in the team became a SQL request aimed at a developer. I led the internal AI contest team as its only engineer, defined the requirements with the people who actually run the queries, and shipped it into their work.
The design rule was that the model never emits SQL. It returns a structured plan; the server renders the SQL; a read-only validation gate decides whether it runs. Failures are sorted into eight types and kept as regression tests.
Result: quote work went from about 40 minutes per case to roughly 4, across ~490 cases a year.
tuto — evidence-tracked video understanding
Python · MIT · Claude Code plugin
The values that matter in a tutorial are often never spoken. They sit on screen, and
auto-captions get them wrong — 1536 becomes 136, and an agent that follows the
caption builds 136-dimensional vectors that fail on every call.
tuto reads the frames. Every knowledge item carries a timestamp, and items grounded in actual pixels are counted separately from items paraphrased off captions. When captions and screen disagree, the screen wins and the conflict is kept as structured data. Merging, cross-checking and the audit line are deterministic Python; exactly one call goes to the model.
I chose the architecture by measuring it rather than guessing. Six configurations on the same 54-minute video: $11.12 → $4.75 (−57%), screen-grounded settings 12 → 40. Benchmark run: 171 knowledge items, 69 screen-grounded, 0% uncited. 785 regression tests. Cost still varies run to run ($4.76–$7.29 on the same benchmark); the cause is traced in the README, along with the limitations I have not solved.
star-coach-analysis — explainable swing coaching
Python · HCI Korea 2026 (equal contribution)
Pose-correction models hand back a fixed skeleton and no reason. STAR-Coach traces the optimization that produced the correction and turns it into an explanation: which body part, at what point in the swing, by how much. In the paper, adding this layer kept PCKh@0.5 at 92.4% on Penn Action with no added processing time.
My part was the explanation layer and the web prototype; the pose-refinement model it sits on is my co-author's prior work. This repo is a mock-data re-implementation of the explanation pipeline, since the paper code lives in a shared account.
Explainable AI-powered Baseball Swing Coaching System Seunghyun Oh, Junsoo Hwang — Proceedings of HCI Korea 2026, pp. 1518–1523, Feb. 2026. Equal contribution.
한국리서치 — Backend Engineer (2024.09 – present), 온라인패널조사부
ASP.NET Core and Web Forms, SQL Server schema and query work, IIS on Windows Server. Legacy code, operational risk, unclear requirements, verification cost. That environment is where the habit came from: before shipping anything a model produced, ask what evidence would make this output checkable.
C# Python Java TypeScript · ASP.NET Core Spring Boot React · MSSQL IIS Nginx
Outside work I keep a small research thread on whether small open LLMs' internal
"being evaluated" signal transfers across languages —
proposal,
proposal stage — no experiments run yet.
Contact · jsjsjsjsjs1019@gmail.com · Portfolio (Notion)

