post-training · evaluation · governance
I research, build, and write about trustworthy AI: model post-training, open evaluation, and the governance around it, focused on high-stakes, regulated domains. This is where that work lives.
A regulation-anchored Trustworthy-AI Scorecard for frontier models in regulated finance. Ten models graded against preregistered, law-derived bars; judgment cases scored by a cross-family LLM panel calibrated to a human rater; every verdict reproducible from frozen transcripts.
Finding: no current model (not Claude Opus 4.8, GPT-5, or Gemini 2.5 Pro) is bare-ready for regulated finance. All ten fail the no-fabrication bar; the value is the spread, and the system wrap each failure demands.
A 2.5B-parameter Turkish language model trained from scratch and post-trained for calibrated honesty: declining what it cannot verify rather than fabricating an answer. Thresholds preregistered before the data existed, scoring by LLM judges validated against hand labels, and every headline number recomputable from the shipped evidence.
Finding: refusal generalized by sentence shape, not by intent. The same harmful request reframed as a thought experiment, a credential, or a roleplay got through. Teaching the intent across eight jailbreak framings, paired with a benign set built on those same framings, took adversarial refusal from 30% to 73% and raised helpfulness instead of costing it.
Interactive: flip a switch and watch a fairness check block a biased change, no code required. An industry-neutral reference for taking an AI solution from data to compliant production: the whole delivery lifecycle as a contract between teams, four stages with a governance spine, and a runnable governance-as-code gate that refuses any change which weakens the system, fairness included. The model is the replaceable part; the contract around it is the product.
Further anchors (healthcare, aviation) and a preference-optimization study are in progress.