Carnegie Mellon · Computer Science

Technical founder.
Builder and researcher.

I build AI, systems, and research prototypes that turn abstract ideas into working technical systems.

Current Astrea, an AI SRE for diagnosing and mitigating production incidents.

An incident reasoning system for production failures. It reconstructs what changed, finds likely root causes, recommends safe mitigations, and evaluates recovery in a realistic incident gym.

Production us east 1 Incident 2481
Astrea production environment during an active incident An interactive service topology connects a deploy to the gateway, API, queue, worker, database, and a proposed rollback. The worker is identified as the likely root cause. gateway worker postgres task queue rollback

Likely root cause

Worker service

Error rate rose seconds after deploy 18f2. Astrea marks this service as the leading root cause candidate.

1 Reconstruct change2 Trace impact3 Test recovery
  • 01 AI + systems
  • 02 Algorithms / optimization
  • 03 Compilers
  • 04 Research prototypes

Systems, models, and the machinery between them.

Five projects, selected for technical depth: how the problem was represented, what was implemented, and how the result was evaluated.

01 / Generative AI

Reasoning-Augmented Diffusion

Generative AI / Structured Reasoning

I built a reasoning-aware text-to-image system that gives diffusion models an explicit intermediate representation before generation. Instead of treating a prompt as one fuzzy embedding, I designed an enriched scene graph that captures objects, counts, attributes, relationships, and reasoning constraints. I also built an image-editing chatbot where the scene graph acts as editable state: when a user asks to “add a window” or “make the couch green,” the graph updates directly and the image regenerates. This made image editing more programmable and controllable than repeatedly prompting a black-box diffusion model.

  • Diffusion Models
  • Scene Graphs
  • LLMs
  • Image Editing
  • Structured Reasoning
  • Python
  • PyTorch
Think Before You Diffuse pipeline showing reasoning, scene graph generation, and diffusion conditioning
System pipelinereasoning → scene graph → image
SGDiff scene-graph-to-image diffusion model overview
Diffusion baselinestructured conditioning
Initial image in chatbot-driven editing pipelineEdited image after the chatbot updates the scene graph
Editable statebefore / after

02 / Compilers

Rust Compiler & Graph Optimization

Compilers / Systems

I built an optimizing L4 compiler in Rust end to end: parsing and elaboration, typechecking with a string interner, lowering to an SSA IR tree, translation to LLVM-like SSA, 3-address x86 SSA, and a backend with dominator analysis, if-conversion, graph-coloring register allocation, SSA deconstruction, and AT&T x86 emission. On the optimization side, I implemented early IR cleanup (copy and constant propagation, folding, DCE), repeated LLVM scalar passes (CSE, arithmetic strength reduction, branch folding), function inlining, tail-call elimination, LICM, partial redundancy elimination, induction-variable strength reduction, scalar replacement, and final x86 peepholes. I tuned the pipeline with per-pass ablations on hot benchmarks. Our O1 pipeline reaches over 70% of optimized GCC.

  • Rust
  • L4
  • Compilers
  • SSA
  • Graph Algorithms
  • x86
  • Optimization
End-to-end L4 compiler pipeline from parsing through optimized x86 output
Compiler architecturefrontend → SSA → x86
Optimization ablation results on the julia L4 benchmark
Pass ablationjulia.l4 benchmark

03 / DL Systems

Accelerating Needle with Graph Fusion

Deep Learning Systems / Compiler Optimizations

We built a minimal PyTorch-like stack entirely on Needle, including tensors and autograd, CPU and CUDA NDArray backends, neural modules (Linear, Conv, activations), optimizers, and CIFAR-10 training with a ResNet-style model. Every layer runs on our own framework rather than PyTorch or TensorFlow. We then implemented TVM-inspired lazy graph compiler passes: constant folding, elementwise chain fusion with a single CUDA interpreter kernel, and matmul, bias, and activation epilogue fusion (tiled GEMM plus activation in registers), plus BroadcastLike so nn.Linear stays lazy. Benchmarks show fewer graph nodes, fewer kernel launches, and faster training versus an unfused baseline.

  • Python
  • CUDA
  • Needle
  • Autograd
  • Graph Optimization
  • Deep Learning Systems
Training time benchmark showing speedup from Needle graph optimization
Training timeoptimized vs. baseline
Benchmark showing fewer graph nodes after optimization
Graph nodesafter optimization
Benchmark showing reduced CUDA kernel launches after graph fusion
Kernel launchesafter fusion

04 / Quantum ML

Qubit_Minds

Quantum ML / Biomedical AI

I built a hybrid quantum-classical classifier for malignant breast cancer diagnosis. Instead of only training a normal neural network, I encoded patient features into a quantum circuit and tested whether that representation could learn a better decision boundary. I compared classical-only, quantum-only, and hybrid PennyLane/PyTorch models, and the hybrid model improved test accuracy by about 2 percentage points over the classical baseline.

  • PennyLane
  • PyTorch
  • Quantum ML
  • Biomedical Classification
  • Hybrid Models
  • Model Evaluation
Quantum variational circuit with angle embedding and entangling layers
Quantum circuitvariational design
Classical neural network model architecture
Classical modelevaluation baseline
Quantum-classical hybrid model architecture
Hybrid modelquantum + classical

05 / Multimodal AI

Lighthouse (TartanHacks 2025)

Multimodal AI / Real-Time Systems

We built Lighthouse, a Chrome extension that captures periodic screen frames through Chrome's offscreen display-media pipeline and streams them to a Flask backend. On each frame, we combine geometry-aware OCR (center- and size-weighted Tesseract bounding boxes via OpenCV) with GPT-4o structured vision outputs such as app or site, activity summary, and salient on-screen objects. A schema-constrained classifier scores distraction against user-defined focus and anti-goals from zero to four, then the extension escalates lock-in notifications and audio nudges in real time. We also built an offline evaluation loop with scraped screenshots and Google's Screen QA dataset to stress-test the pipeline before demoing live.

  • Python
  • Chrome Extension
  • Flask
  • Computer Vision
  • GPT-4o
  • OCR
  • Real-Time Systems
Lighthouse system architecture from Chrome capture through OCR and multimodal language models
System architecturecapture → classify → intervene
Chrome browser showing the Lighthouse extension focus and distraction goals
Extension UIfocus + anti-goals
Lighthouse detecting distractions and sending lock-in notifications
Real-time responsedetection + nudge

I care about the part between an interesting idea and a result you can trust.

I'm at CMU studying CS. I tend to get stuck on the unglamorous part: how you represent the problem, how you measure whether it is actually working, and whether the pipeline is testing your real hypothesis or just producing a convincing demo.

The work above is where that obsession showed up.

Working on hard infrastructure, AI systems, or technical research?

I’m always interested in comparing notes with people building close to the machinery.