Picture for Luo Mai

Luo Mai

BatchGen: An Architecture for Scalable and Efficient Batch Inference

Add code
Jun 19, 2026
Viaarxiv icon

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems

Add code
Jun 19, 2026
Viaarxiv icon

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

Add code
May 30, 2026
Viaarxiv icon

Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents

Add code
May 27, 2026
Viaarxiv icon

RAGBoost: Efficient Retrieval-Augmented Generation with Accuracy-Preserving Context Reuse

Add code
Nov 05, 2025
Viaarxiv icon

HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing

Add code
May 18, 2025
Figure 1 for HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
Figure 2 for HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
Figure 3 for HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
Figure 4 for HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
Viaarxiv icon

MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems

Add code
May 16, 2025
Viaarxiv icon

MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching

Add code
Mar 12, 2025
Viaarxiv icon

WaferLLM: A Wafer-Scale LLM Inference System

Add code
Feb 06, 2025
Viaarxiv icon

MoE-CAP: Cost-Accuracy-Performance Benchmarking for Mixture-of-Experts Systems

Add code
Dec 10, 2024
Viaarxiv icon