Evaluation Open Source Projects
Browse 587 Evaluation open source projects, ranked by GitHub stars. Find the most popular Evaluation tools and libraries.
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Metrics details
| Stars | 23,420 |
NVIDIA/Megatron-LM
Ongoing research training transformer models at scale
Metrics details
| Stars | 17,110 |
EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models.
Metrics details
| Stars | 13,340 |
mrgloom/awesome-semantic-segmentation
:metal: awesome-semantic-segmentation
Metrics details
| Stars | 10,849 |
open-compass/opencompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Metrics details
| Stars | 7,209 |
PaddlePaddle/PaddleClas
A treasure chest for visual classification and recognition powered by PaddlePaddle
Metrics details
| Stars | 5,823 |
rhaiscript/rhai
Rhai - An embedded scripting language for Rust.
Metrics details
| Stars | 5,490 |
remeda/remeda
A utility library for JavaScript and TypeScript.
Metrics details
| Stars | 5,398 |
nianticlabs/monodepth2
[ICCV 2019] Monocular depth estimation from a single image
Metrics details
| Stars | 4,499 |
open-compass/VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Metrics details
| Stars | 4,291 |
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Metrics details
| Stars | 4,281 |
deepmind/learning-to-learn
Learning to Learn in TensorFlow
Metrics details
| Stars | 4,066 |
Knetic/govaluate
Arbitrary expression evaluation for golang
Metrics details
| Stars | 3,940 |
clovaai/deep-text-recognition-benchmark
Text recognition (optical character recognition) with deep learning methods, ICCV 2019
Metrics details
| Stars | 3,937 |
sdiehl/write-you-a-haskell
Building a modern functional compiler from first principles. (http://dev.stephendiehl.com/fun/)
Metrics details
| Stars | 3,484 |
CLUEbenchmark/SuperCLUE
SuperCLUE: 中文通用大模型综合性基准 | A Benchmark for Foundation Models in Chinese
Metrics details
| Stars | 3,296 |
thomasahle/sunfish
Sunfish: a Python Chess Engine in 111 lines of code
Metrics details
| Stars | 3,261 |
viebel/klipse
Klipse is a JavaScript plugin for embedding interactive code snippets in tech blogs.
Metrics details
| Stars | 3,134 |
modelscope/llmuses
A streamlined and customizable framework for efficient large model evaluation and performance benchmarking
Metrics details
| Stars | 3,086 |
ianarawjo/ChainForge
An open-source visual programming environment for battle-testing prompts to LLMs.
Metrics details
| Stars | 3,013 |
google/cel-go
Fast, portable, non-Turing complete expression evaluation with gradual typing (Go)
Metrics details
| Stars | 3,000 |
zzw922cn/Automatic_Speech_Recognition
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
Metrics details
| Stars | 2,834 |
microsoft/promptbench
A unified evaluation framework for large language models
Metrics details
| Stars | 2,817 |
microsoft/CodeBERT
CodeBERT
Metrics details
| Stars | 2,786 |
nix-community/nix-direnv
A fast, persistent use_nix/use_flake implementation for direnv [maintainer=@Mic92 / @bbenne10]
Metrics details
| Stars | 2,733 |
Project-MONAI/tutorials
MONAI Tutorials
Metrics details
| Stars | 2,513 |
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Metrics details
| Stars | 2,466 |
junfu1115/DANet
Dual Attention Network for Scene Segmentation (CVPR2019)
Metrics details
| Stars | 2,460 |
Zhongdao/Towards-Realtime-MOT
Joint Detection and Embedding for fast multi-object tracking
Metrics details
| Stars | 2,441 |
neuecc/ZeroFormatter
Infinitely Fast Deserializer for .NET, .NET Core and Unity.
Metrics details
| Stars | 2,426 |
uptrain-ai/uptrain
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
Metrics details
| Stars | 2,355 |
Xnhyacinth/Awesome-LLM-Long-Context-Modeling
📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
Metrics details
| Stars | 2,146 |
bluekitchen/btstack
Dual-mode Bluetooth stack, with small memory footprint.
Metrics details
| Stars | 2,115 |
sfujim/TD3
Author's PyTorch implementation of TD3 for OpenAI gym tasks
Metrics details
| Stars | 2,095 |
openai/image-gpt
Metrics details
| Stars | 2,094 |
ContinualAI/avalanche
Avalanche: an End-to-End Library for Continual Learning based on PyTorch.
Metrics details
| Stars | 2,072 |
Cloud-CV/EvalAI
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI
Metrics details
| Stars | 2,032 |
tinghuiz/SfMLearner
An unsupervised learning framework for depth and ego-motion estimation from monocular videos
Metrics details
| Stars | 2,014 |
tatsu-lab/alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
Metrics details
| Stars | 2,004 |
nix-community/NUR
Nix User Repository: User contributed nix packages [maintainer=@Pandapip1]
Metrics details
| Stars | 1,901 |
magicleap/Atlas
Atlas: End-to-End 3D Scene Reconstruction from Posed Images
Metrics details
| Stars | 1,857 |
xinshuoweng/AB3DMOT
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
Metrics details
| Stars | 1,842 |
nywang16/Pixel2Mesh
Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. In ECCV2018.
Metrics details
| Stars | 1,760 |
autonomousvision/occupancy_networks
This repository contains the code for the paper "Occupancy Networks - Learning 3D Reconstruction in Function Space"
Metrics details
| Stars | 1,660 |
hszhao/PSPNet
Pyramid Scene Parsing Network, CVPR2017.
Metrics details
| Stars | 1,659 |
google-research/pegasus
Metrics details
| Stars | 1,657 |
benhamner/Metrics
Machine learning evaluation metrics, implemented in Python, R, Haskell, and MATLAB / Octave
Metrics details
| Stars | 1,650 |
timoschick/pet
This repository contains the code for "Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference"
Metrics details
| Stars | 1,625 |
InterDigitalInc/CompressAI
A PyTorch library and evaluation platform for end-to-end compression research
Metrics details
| Stars | 1,612 |
experiencor/keras-yolo3
Training and Detecting Objects with YOLO3
Metrics details
| Stars | 1,608 |
MLGroupJLU/LLM-eval-survey
The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".
Metrics details
| Stars | 1,607 |
jcjohnson/densecap
Dense image captioning in Torch
Metrics details
| Stars | 1,598 |
dennybritz/chatbot-retrieval
Dual LSTM Encoder for Dialog Response Generation
Metrics details
| Stars | 1,570 |
sepandhaghighi/pycm
Multi-class confusion matrix library in Python
Metrics details
| Stars | 1,506 |
Maluuba/nlg-eval
Evaluation code for various unsupervised automated metrics for Natural Language Generation.
Metrics details
| Stars | 1,391 |
natasha/natasha
Solves basic Russian NLP tasks, API for lower level Natasha projects
Metrics details
| Stars | 1,342 |
NVlabs/DG-Net
:couple: Joint Discriminative and Generative Learning for Person Re-identification. CVPR'19 (Oral) :couple:
Metrics details
| Stars | 1,299 |
abo-abo/lispy
Short and sweet LISP editing
Metrics details
| Stars | 1,297 |
j96w/DenseFusion
"DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion" code repository
Metrics details
| Stars | 1,273 |
EthicalML/xai
XAI - An eXplainability toolbox for machine learning
Metrics details
| Stars | 1,254 |
