Emory University
PhD, Computer Science and InformaticsAdvised by Prof. Jinho D. Choi, EmoryNLP Lab.
PhD Candidate, Computer Science & Informatics — Emory University
My research asks how language models fail when it matters most — then builds the benchmarks, safety evaluations, and retrieval systems to catch it first.
I'm a PhD candidate in Computer Science and Informatics at Emory University, working with Prof. Jinho D. Choi in the EmoryNLP Lab. My research sits at the intersection of language and safety: constructing clinician-annotated benchmarks for crisis detection, studying failure modes like sycophancy in multi-turn dialogue, and building retrieval-augmented systems that hold up against adversarial and proprietary document collections. I interned on the Search and Embeddings team at Cohere, working on context-pruning rerankers.
Before Emory, I completed a master's in Computational Linguistics at Seoul National University, where my dissertation focused on simplifying LLM alignment and detoxification through instruction and preference data, and a BA in Linguistics at Korea University.
Advised by Prof. Jinho D. Choi, EmoryNLP Lab.
Dissertation: Simplifying Large Language Model Alignment and Detoxification — Comprehensive Instruction and Preference Data Solutions.
Built a generative AI assistant that analyzes crash test data to support engineering decisions for vehicle safety.
Developed a Korean-specialized multilingual LLM using Mixture of Experts, tokenizer expansion, instruction tuning, DPO, and ORPO; deployed publicly at dag.snu.ac.kr.
Developed a domain-specific LLM using financial raw data provided by South Korea's central bank.
Led compilation of a decade of medical and legal resources, working directly with faculty from KU's Medical and Law schools.
Ran weekly Python training for a group of 10 students over a full semester.
Secure Multifaceted-RAG for Enterprise: Hybrid Knowledge Retrieval with Security Filtering
Reference-Aligned Retrieval-Augmented Question Answering over Heterogeneous Proprietary Documents
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Models
Large Language Model Detoxification: Data and Metric Solutions
The Manchu Milestone: Wav2vec 2.0 and the Future of Minority Language Speech Recognition
Korean Bio-Medical Corpus (KBMC) for Named Entity Recognition with Span Representation Analysis
Mergen: The First Manchu–Korean Machine Translation Model Trained on Augmented Data
Teaching assistant for Computational Linguistics (CS329), Research Practicum in Artificial Intelligence (CS371), and Introduction to Artificial Intelligence (CS211).
Delivered lectures on LLM training methods and practical use of the ChatGPT API to 20 undergraduates.
Assisted "Basics of Natural Language Processing"; earned a distinguished academic scholarship for the role.
Graded assignments and developed course textbooks.