Curriculum Vitae

- Instruction Tuning (SFT) & Domain Alignment: Designed a multi-turn complex SFT pipeline for vertical dialogue and task planning; engineered prompt templates and preference datasets for open-source foundation models (Qwen/Llama) using LLaMA-Factory, mitigating hallucination and boosting benchmark accuracy by +23%.
- Preference Alignment (DPO/RLHF): Implemented preference alignment with LLM-as-a-Judge multidimensional scoring and fine-grained ranking, significantly enhancing communication quality, safety compliance, and multi-turn logical consistency.
- Distributed Training (Slurm & ZeRO-3): Deployed multi-node multi-GPU clusters (A100/H100) via Slurm on LRZ HPC; integrated DeepSpeed ZeRO-3, FlashAttention-2, and gradient accumulation optimization, reducing per-GPU memory usage from 78GB to 52GB (-33%) and accelerating end-to-end throughput by +28%.
- High-Concurrency Serving (vLLM): Built low-latency, high-throughput model serving pipelines with vLLM; leveraged PagedAttention, Continuous Batching, and Prefix Caching to optimize KV cache reuse, reducing long-context Time-To-First-Token (TTFT) by -45%.
- Embodied Agent & Spatial Grounding: Researched 3D-LLM multimodal feature alignment using Cross-Attention and geometric embeddings (FPFH/PointNet++), resolving physical interaction ambiguities in complex spatial instruction following.
- Multimodal Geometric Fusion & World Model: Combined RANSAC robust estimation, Gauss-Newton optimization, and ICP/NDT registration to counter occlusion and point cloud noise, cutting pose estimation error by -65% (RMSE 5.2mm → 1.8mm) and single-frame latency by -62.5% (120ms → 45ms) for upstream agent planning.
- Full-Duplex Conversational Loop: Engineered an end-to-end streaming conversational pipeline (Streaming ASR → LLM Intent Reasoning → Streaming TTS) for natural, fluid anthropomorphic customer service dialogue.
- Context Routing & Barge-in Handling: Developed sliding-window dynamic context stitching and intent routing modules, robustly maintaining long-horizon context consistency with seamless user interruption (barge-in) flow control.
- Sub-Second Latency Optimization: Re-architected chunked transport with asyncio, WebSockets, Ring Buffers, and vLLM Chunked Prefill with INT8 quantization, slashing first-frame voice response latency from 3.4s to <0.9s.
- Agentic Workflow & Tool Use: Built an autonomous code repair agent on Gemini featuring closed-loop localization, patch planning, sandboxed execution, and self-reflection (100% tool invocation success, 11.40% end-to-end task success in complex long-context environments).
- Benchmark Evaluation & SOTA Performance: Evaluated against open-source SOTA baselines (Agentless, OpenHands) on automated benchmarks, achieving execution verification metrics of EV-Micro 8.50% and EV-Macro 23.03%.
- High-Availability Sandboxed System: Built an asynchronous orchestration hub with Django REST Framework, Celery, Redis, and containerized Docker sandboxes with automated Pytest feedback loops, averaging only ~0.35M tokens per task.
Proficient in Python for machine learning workflows and systems programming; experienced with Linux HPC clusters, shell scripting, containerized environments, and Git-based collaborative development.
Deep understanding of Transformer architectures; hands-on fine-tuning and alignment of open-source LLMs (Qwen/Llama); automated multidimensional quality evaluation and hallucination mitigation.
Proficient in designing autonomous agent closed loops with self-reflection; full-duplex streaming conversational pipelines; multimodal spatial instruction following and 3D geometric grounding.
Multi-node multi-GPU cluster distributed training; high-throughput low-latency inference serving; asynchronous architecture design with containerized sandbox isolation.