>_Skillful
Need help with advanced AI agent engineering?Contact FirmAdapt

Search

FasterTransformer

NVIDIA Framework for LLM Inference(Transitioned to TensorRT-LLM)

AgentLLM Inference
6.4K1 dir

MInference

To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

...more
AgentLLM Inference
1.2K1 dir

exllama

A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.

AgentLLM Inference
2.9K1 dir

mistral.rs

Blazingly fast LLM inference.

AgentLLM Inference
6.7K1 dir

DeepSpeed-Mii

MII makes low-latency and high-throughput inference, similar to vLLM powered by DeepSpeed.

AgentLLM Inference
2.1K1 dir

Text-Embeddings-Inference

Inference for text-embeddings in Rust, HFOIL Licence.

AgentLLM Inference
4.6K1 dir

Infinity

Inference for text-embeddings in Python

AgentLLM Inference
2.7K1 dir

prima.cpp

A distributed implementation of llama.cpp that lets you run 70B-level LLMs on your everyday devices.

AgentLLM Inference
1 dir

deploy-llms-with-ansible

Easily deploy any LLM on a VM with minimal configuration, using Ansible.

AgentLLM Inference
31 dir

wechat-chatgpt

Use ChatGPT On Wechat via wechaty

AgentLLM Applications
13K1 dir

Serge

a chat interface crafted with llama.cpp for running Alpaca models. No API keys, entirely self-hosted!

AgentLLM Applications
5.7K1 dir

IntelliServer

simplifies the evaluation of LLMs by providing a unified microservice to access and test multiple AI models.

AgentLLM Applications
291 dir

Search with Lepton

Build your own conversational search engine using less than 500 lines of code by [LeptonAI](https://github.com/leptonai).

...more
AgentLLM Applications
8.1K1 dir

Tune Studio

Playground for devs to finetune & deploy LLMs

AgentLLM Applications
1 dir

talkd.ai dialog

Simple API for deploying any RAG or LLM that you want adding plugins.

AgentLLM Applications
4311 dir

Wllama

WebAssembly binding for llama.cpp - Enabling in-browser LLM inference

AgentLLM Applications
1K1 dir

GPUStack

An open-source GPU cluster manager for running LLMs

AgentLLM Applications
4.7K1 dir

MNN-LLM

- A Device-Inference framework, including LLM Inference on device(Mobile Phone/PC/IOT)

AgentLLM Applications
15K1 dir

CAMEL

First LLM Multi-agent framework.

AgentLLM Applications
1 dir

QA-Pilot

An interactive chat project that leverages Ollama/OpenAI/MistralAI LLMs for rapid understanding and navigation of GitHub code repository or compressed file resources.

...more
AgentLLM Applications
3161 dir