Evaluation Open Source Projects

Browse 587 Evaluation open source projects, ranked by GitHub stars. Find the most popular Evaluation tools and libraries.

Share your experience:✍️ Write a Post❓ Ask a Question
1-60 of 587 projects
23,420 stars

promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

Metrics details
Stars23,420
17,110 stars

NVIDIA/Megatron-LM

Ongoing research training transformer models at scale

Metrics details
Stars17,110
13,340 stars

EleutherAI/lm-evaluation-harness

A framework for few-shot evaluation of language models.

Metrics details
Stars13,340
10,849 stars

mrgloom/awesome-semantic-segmentation

:metal: awesome-semantic-segmentation

Metrics details
Stars10,849
7,209 stars

open-compass/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

Metrics details
Stars7,209
5,823 stars

PaddlePaddle/PaddleClas

A treasure chest for visual classification and recognition powered by PaddlePaddle

Metrics details
Stars5,823
5,490 stars

rhaiscript/rhai

Rhai - An embedded scripting language for Rust.

Metrics details
Stars5,490
5,398 stars

remeda/remeda

A utility library for JavaScript and TypeScript.

Metrics details
Stars5,398
4,499 stars

nianticlabs/monodepth2

[ICCV 2019] Monocular depth estimation from a single image

Metrics details
Stars4,499
4,291 stars

open-compass/VLMEvalKit

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

Metrics details
Stars4,291
4,281 stars

MichaelGrupp/evo

Python package for the evaluation of odometry and SLAM

Metrics details
Stars4,281
4,066 stars

deepmind/learning-to-learn

Learning to Learn in TensorFlow

Metrics details
Stars4,066
3,940 stars

Knetic/govaluate

Arbitrary expression evaluation for golang

Metrics details
Stars3,940
3,937 stars

clovaai/deep-text-recognition-benchmark

Text recognition (optical character recognition) with deep learning methods, ICCV 2019

Metrics details
Stars3,937
3,484 stars

sdiehl/write-you-a-haskell

Building a modern functional compiler from first principles. (http://dev.stephendiehl.com/fun/)

Metrics details
Stars3,484
3,296 stars

CLUEbenchmark/SuperCLUE

SuperCLUE: 中文通用大模型综合性基准 | A Benchmark for Foundation Models in Chinese

Metrics details
Stars3,296
3,261 stars

thomasahle/sunfish

Sunfish: a Python Chess Engine in 111 lines of code

Metrics details
Stars3,261
3,134 stars

viebel/klipse

Klipse is a JavaScript plugin for embedding interactive code snippets in tech blogs.

Metrics details
Stars3,134
3,086 stars

modelscope/llmuses

A streamlined and customizable framework for efficient large model evaluation and performance benchmarking

Metrics details
Stars3,086
3,013 stars

ianarawjo/ChainForge

An open-source visual programming environment for battle-testing prompts to LLMs.

Metrics details
Stars3,013
3,000 stars

google/cel-go

Fast, portable, non-Turing complete expression evaluation with gradual typing (Go)

Metrics details
Stars3,000
2,834 stars

zzw922cn/Automatic_Speech_Recognition

End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow

Metrics details
Stars2,834
2,817 stars

microsoft/promptbench

A unified evaluation framework for large language models

Metrics details
Stars2,817
2,786 stars

microsoft/CodeBERT

CodeBERT

Metrics details
Stars2,786
2,733 stars

nix-community/nix-direnv

A fast, persistent use_nix/use_flake implementation for direnv [maintainer=@Mic92 / @bbenne10]

Metrics details
Stars2,733
2,513 stars

Project-MONAI/tutorials

MONAI Tutorials

Metrics details
Stars2,513
2,466 stars

huggingface/evaluate

🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

Metrics details
Stars2,466
2,460 stars

junfu1115/DANet

Dual Attention Network for Scene Segmentation (CVPR2019)

Metrics details
Stars2,460
2,441 stars

Zhongdao/Towards-Realtime-MOT

Joint Detection and Embedding for fast multi-object tracking

Metrics details
Stars2,441
2,426 stars

neuecc/ZeroFormatter

Infinitely Fast Deserializer for .NET, .NET Core and Unity.

Metrics details
Stars2,426
2,355 stars

uptrain-ai/uptrain

UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.

Metrics details
Stars2,355
2,146 stars

Xnhyacinth/Awesome-LLM-Long-Context-Modeling

📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥

Metrics details
Stars2,146
2,115 stars

bluekitchen/btstack

Dual-mode Bluetooth stack, with small memory footprint.

Metrics details
Stars2,115
2,095 stars

sfujim/TD3

Author's PyTorch implementation of TD3 for OpenAI gym tasks

Metrics details
Stars2,095
2,094 stars

openai/image-gpt

Metrics details
Stars2,094
2,072 stars

ContinualAI/avalanche

Avalanche: an End-to-End Library for Continual Learning based on PyTorch.

Metrics details
Stars2,072
2,032 stars

Cloud-CV/EvalAI

:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI

Metrics details
Stars2,032
2,014 stars

tinghuiz/SfMLearner

An unsupervised learning framework for depth and ego-motion estimation from monocular videos

Metrics details
Stars2,014
2,004 stars

tatsu-lab/alpaca_eval

An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.

Metrics details
Stars2,004
1,901 stars

nix-community/NUR

Nix User Repository: User contributed nix packages [maintainer=@Pandapip1]

Metrics details
Stars1,901
1,857 stars

magicleap/Atlas

Atlas: End-to-End 3D Scene Reconstruction from Posed Images

Metrics details
Stars1,857
1,842 stars

xinshuoweng/AB3DMOT

(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

Metrics details
Stars1,842
1,760 stars

nywang16/Pixel2Mesh

Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. In ECCV2018.

Metrics details
Stars1,760
1,660 stars

autonomousvision/occupancy_networks

This repository contains the code for the paper "Occupancy Networks - Learning 3D Reconstruction in Function Space"

Metrics details
Stars1,660
1,659 stars

hszhao/PSPNet

Pyramid Scene Parsing Network, CVPR2017.

Metrics details
Stars1,659
1,657 stars

google-research/pegasus

Metrics details
Stars1,657
1,650 stars

benhamner/Metrics

Machine learning evaluation metrics, implemented in Python, R, Haskell, and MATLAB / Octave

Metrics details
Stars1,650
1,625 stars

timoschick/pet

This repository contains the code for "Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference"

Metrics details
Stars1,625
1,612 stars

InterDigitalInc/CompressAI

A PyTorch library and evaluation platform for end-to-end compression research

Metrics details
Stars1,612
1,608 stars

experiencor/keras-yolo3

Training and Detecting Objects with YOLO3

Metrics details
Stars1,608
1,607 stars

MLGroupJLU/LLM-eval-survey

The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

Metrics details
Stars1,607
1,598 stars

jcjohnson/densecap

Dense image captioning in Torch

Metrics details
Stars1,598
1,570 stars

dennybritz/chatbot-retrieval

Dual LSTM Encoder for Dialog Response Generation

Metrics details
Stars1,570
1,506 stars

sepandhaghighi/pycm

Multi-class confusion matrix library in Python

Metrics details
Stars1,506
1,391 stars

Maluuba/nlg-eval

Evaluation code for various unsupervised automated metrics for Natural Language Generation.

Metrics details
Stars1,391
1,342 stars

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

Metrics details
Stars1,342
1,299 stars

NVlabs/DG-Net

:couple: Joint Discriminative and Generative Learning for Person Re-identification. CVPR'19 (Oral) :couple:

Metrics details
Stars1,299
1,297 stars

abo-abo/lispy

Short and sweet LISP editing

Metrics details
Stars1,297
1,273 stars

j96w/DenseFusion

"DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion" code repository

Metrics details
Stars1,273
1,254 stars

EthicalML/xai

XAI - An eXplainability toolbox for machine learning

Metrics details
Stars1,254
1-60 of 587 projects
Get A Weekly Email With Trending Evaluation Projects
Stay updated on Evaluation plus related topics you pick below.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.