inference-serving
Popular repositories Loading
-
energy-inference
energy-inference PublicForked from grantwilkins/energy-inference
Code for MPhil Thesis at University of Cambridge
Python 1
-
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
-
Sequence-Scheduling
Sequence-Scheduling PublicForked from zhengzangw/Sequence-Scheduling
PyTorch implementation of paper "Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline".
Python
-
LLM-serving-with-proxy-models
LLM-serving-with-proxy-models PublicForked from James-QiuHaoran/LLM-serving-with-proxy-models
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Jupyter Notebook
-
pytorch-transformer
pytorch-transformer PublicForked from hkproj/pytorch-transformer
Attention is all you need implementation
Jupyter Notebook
Repositories
- Llamacpp-Latency-Profiler Public Forked from leexcape/Llamacpp-Latency-Profiler
A lightweight benchmarking tool for llama.cpp based transformer models. Measures generation latency, tokens per second, and captures device + model metadata across diverse hardware (CPU, GPU, Jetson, Raspberry Pi, etc.). Results are saved in timestamped JSON logs for easy comparison and analysis.
- LLM-Latency-Profiler Public Forked from leexcape/LLM-Latency-Profiler
A lightweight benchmarking tool for Hugging Face LLMs using AutoModelForCausalLM. Measures generation latency, tokens per second, and captures device + model metadata across diverse hardware (CPU, GPU, Jetson, Raspberry Pi, etc.). Results are saved in timestamped JSON logs for easy comparison and analysis.
- infrastructure Public
- orion Public Forked from eth-easl/orion
An interference-aware scheduler for fine-grained GPU sharing
- energy-inference Public Forked from grantwilkins/energy-inference
Code for MPhil Thesis at University of Cambridge
- ni-science-festival Public
- TAO-Amodal Public Forked from WesleyHsieh0806/TAO-Amodal
Official Code for Tracking Any Object Amodally
- LLM-serving-with-proxy-models Public Forked from James-QiuHaoran/LLM-serving-with-proxy-models
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
- deepstream_python_apps Public Forked from NVIDIA-AI-IOT/deepstream_python_apps
DeepStream SDK Python bindings and sample applications
- Programming-Massively-Parallel-Processors Public Forked from R100001/Programming-Massively-Parallel-Processors
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…