FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge Devices
First author (1/5)
Adaptive DNN inference under tight device-memory budgets.
- Memory Management
- Scheduling
- Adaptive Inference
Edge AI
Locate on map ↗Systems · Models · Physical intelligence
I build efficient AI systems for resource-constrained devices, and explore how these systems enable embodied models and self-evolving physical intelligence.
Full paper list →First author (1/5)
Adaptive DNN inference under tight device-memory budgets.
Edge AI
Locate on map ↗Co-first author (1/6)
Vector table lookup makes parallel ultra-low-bit LLM inference fast on edge CPUs.
Large Language Models (LLMs)Edge AI
Locate on map ↗Third author (3/6)
Dynamic width selection on nested models for efficient on-device LLM decoding.
Large Language Models (LLMs)Edge AI
Locate on map ↗Fifth author (5/8)
Cross-model scheduling for efficient multi-DNN edge video analytics.
Edge AIVideo Analytics
Locate on map ↗Sixth author (6/9)
Pruning-efficient, parallel graph mining with near-memory computing.
Sole author (1/1)
A synthesis of inference systems for resource-constrained edge AI deployment.
Edge AI
Locate on map ↗First author (1/6)
One shared KV cache coordinates concurrent action and language generation in VLAs.
Vision-Language-Action Models (VLAs)Edge AI
Locate on map ↗Fifth author (5/11)
A portable inference runtime for embodied models across heterogeneous robots.
Vision-Language-Action Models (VLAs)World-Action Models (WAMs)Edge AI
Locate on map ↗Co-first author (2/11)
Action-space probes detect failures early in generative robot policies.
Vision-Language-Action Models (VLAs)
Locate on map ↗Fifth author (5/11)
In-context causal learning for generalizable embodied manipulation.
Vision-Language-Action Models (VLAs)World-Action Models (WAMs)
Locate on map ↗Seventh author (7/15)
A closed-loop embodied harness for self-evolving physical intelligence.
Embodied Agents
Locate on map ↗Ninth author (9/15)
Skill-aware reflection helps embodied agents evolve through experience.
Embodied Agents
Locate on map ↗Fourth author · efficiency section lead (4/25)
A survey of personal LLM agents, with a focus on capability, efficiency and security.
Personal AgentsLarge Language Models (LLMs)
Locate on map ↗Sixth author (6/9)
An LLM-based framework that unifies synthetic sensing.
Large Language Models (LLMs)
Locate on map ↗Sixth author (6/11)
An empirical study of LLM reasoning with strict output-length budgets.
Large Language Models (LLMs)Edge AI
Locate on map ↗Project maintainer (project)
Quantization, kernels and parallel rollouts bring world-action models to 24 GB GPUs.
World-Action Models (WAMs)
Locate on map ↗Open-source contributor (contribution)
Correct zero-point handling for low-bit and asymmetric GPTQv2 inference in vLLM.
Large Language Models (LLMs)
Locate on map ↗No matching works. Try another search or reset the filters.