Benchmark LLMs With More Consistency

LM Evaluation Harness helps teams run repeatable few-shot evaluations when comparing language models and tracking quality.

26
Hotness score
49
Reliability score
13,340
Stars
almost 6 years
Age
0
Published reviews
0
Questions

Scores

Hotness and reliability

The two headline scores, measured daily, with 30, 90, and 180 day views.

Hotness score

Jul 19, 2026

Current26

Previous27

Weekly average

100500
Full metrics details

Derived only from GitHub GraphQL starredAt events after daily star totals are reconciled.

180 observed daily rows. Missing days are not fabricated.

Hotness formula

40% Hot today + 40% Hot this week + 20% Breakout.

  • 40% Hot today: 14
  • 40% Hot this week: 37
  • 20% Breakout: 1
  • Stars gained: 1d: 3
  • Stars gained: 7d: 67
  • Stars gained: 14d: 139
  • Stars gained: 30d: 325
  • Stars gained: 90d: 13340

Per-day formula: 0.40 × Hot today + 0.40 × Hot week + 0.20 × Breakout.

  • Same-day multiplier = stars 1d ÷ max(1, stars 7d ÷ 7)
  • Weekly multiplier = stars 7d ÷ max(1, stars 30d ÷ 30 × 7)
  • Fortnight multiplier = stars 14d ÷ max(1, stars 90d ÷ 90 × 14)
  • Hot today = clamp(70 × log-scale(stars 1d, 1000) + 30 × breakout-scale(same-day, 4))
  • Hot week = clamp(75 × log-scale(stars 7d, 5000) + 25 × breakout-scale(weekly, 3))
  • Breakout = clamp(35 × breakout-scale(same-day, 4) + 40 × breakout-scale(weekly, 3) + 25 × breakout-scale(fortnight, 2.5))
DateStarsStars 1dStars 7dStars 14dStars 30dStars 90dSame-day multiplierWeekly multiplierFortnight multiplierHot todayHot weekBreakoutStored score

Reliability score

Jul 19, 2026

Current49

Previous49

Weekly average

100500
Full metrics details

Derived from the displayed continuity, closure, shipping, liveness, and support-burden components.

180 observed daily rows. Missing days are not fabricated.

Reliability formula

30% Continuity + 30% Closure + 20% Shipping + 10% Liveness + 10% Support burden.

  • 30% Continuity: 67
  • 30% Closure: 48
  • 20% Shipping: 18
  • 10% Liveness: 100
  • 10% Support burden: 14

Per-day formula: 0.30 × Continuity + 0.30 × Closure + 0.20 × Shipping + 0.10 × Liveness + 0.10 × Support burden.

  • Continuity = 35% active ratio 90d + 20% active ratio 30d + 15% PR efficiency 90d + 10% PR efficiency 30d + 10% liveness + 10% observed-history coverage
  • Closure = 55% PR merge efficiency 90d + 45% issue close efficiency 90d
  • Shipping = 100 × (20% × release ratio 30d + 30% × release ratio 90d + 50% × release ratio 180d); ratios are releases ÷ 2, 6, and 12, capped at 1
  • Support burden = 100 − 2 × open issues per 1,000 stars
  • Liveness = 100 × exp(−pushed days ago ÷ 120)

Stars, issues, pull requests, and releases are daily observations. Pushed-days and license are repository snapshot-derived inputs and are not independent historical GitHub events.

DateIssues openedIssues closedPRs openedPRs mergedReleasesStarsOpen issuesPushed days agoContinuityClosureShippingLivenessSupport burdenLicenseObserved daysMissing daysNeeds healingStored score

Reliability breakdown

What drives the reliability score

Component scores measured daily from GitHub activity (2026-07-18).

Continuity30% of headline67

Activity on 1 of 90 tracked days in the last 90 and 0 of 30 in the last 30.

Closure30% of headline48

Merged 0 of 0 PRs opened and closed 0 of 0 issues opened over the last 90 tracked days — PR flow carries 55% of this component, issue flow 45%.

Shipping20% of headline18

3 releases in the last 180 days, 1 in the last 90 and 0 in the last 30 — a steady cadence scores highest.

Liveness10% of headline100

Last push 7 days ago — freshness decays as pushes age (roughly halves every 83 days without a push).

Support burden10% of headline14

575 open issues against 13,340 stars — about 43 open issues per 1,000 stars. Lighter backlogs score higher.

Adoption confidence67

Modeled — how confidently teams are adopting this repo. Stargazers

Maintenance quality51

Modeled from PR/issue responsiveness and upkeep signals. Activity pulse

Risk score52

Modeled — lower is better; adoption and continuity risk. Repository

Stays active (30d)63%

Modeled chance the repo stays active over 30 days. Activity pulse

Stays active (90d)66%

Modeled chance the repo stays active over 90 days. Activity pulse

Release rhythm (180d)18

Regularity of releases over the last 180 days. Releases

Maintainer bus risk (90d)31%

Modeled — lower is better; concentration of commits in few maintainers. Contributors graph

Topics

Explore related topics

Jump into the topic listings this repository belongs to.

Measured history

Project metrics

Weekly GitHub totals with 30, 90, and 180 day views.

Issues opened

Mar 8, 2026

Current5

Previous7

Weekly total

740
Full metrics details

GitHub API observation. Historical values render only when the source provides a real observation for that day.

46 observed daily rows. Missing days are not fabricated.

DateValue

Issues closed

Mar 8, 2026

Current4

Previous3

Weekly total

420
Full metrics details

GitHub Search issue totals queried for exact UTC days; rolling values are sums of proven daily counts.

46 observed daily rows. Missing days are not fabricated.

DateValue

Pull requests opened

Mar 8, 2026

Current3

Previous10

Weekly total

21110
Full metrics details

GitHub API observation. Historical values render only when the source provides a real observation for that day.

46 observed daily rows. Missing days are not fabricated.

DateValue

Pull requests closed

Mar 8, 2026

Current4

Previous3

Weekly total

1680
Full metrics details

GitHub Search pull-request totals queried for exact UTC days; rolling values are sums of proven daily counts.

46 observed daily rows. Missing days are not fabricated.

DateValue

Pull requests merged

Mar 8, 2026

Current4

Previous1

Weekly total

1370
Full metrics details

GitHub API observation. Historical values render only when the source provides a real observation for that day.

46 observed daily rows. Missing days are not fabricated.

DateValue

Issue close ratio (daily)

Mar 6, 2026

Current0.00

Previous4.00

Daily ratio — a gap means nothing was opened that day

How this works: each day's value is issues closed that day ÷ issues opened that day. Days when nothing was opened render as gaps — never a rolling average or a made-up zero. Values above 1 mean the project closed more issues than arrived that day.

4.002.000.00
Full metrics details

GitHub Search issue totals queried for exact UTC days; rolling values are sums of proven daily counts.

19 observed daily rows. Missing days are not fabricated.

Per-day formula: issues closed that day ÷ issues opened that day; blank when nothing was opened (a gap, not a zero).

DateIssues openedIssues closedStored score

Pull request close ratio (daily)

Mar 5, 2026

Current0.50

Previous3.00

Daily ratio — a gap means nothing was opened that day

How this works: each day's value is pull requests closed that day ÷ pull requests opened that day. Days when nothing was opened render as gaps — never a rolling average or a made-up zero. Values above 1 mean the project closed more pull requests than arrived that day.

3.001.500.00
Full metrics details

GitHub Search pull-request totals queried for exact UTC days; rolling values are sums of proven daily counts.

34 observed daily rows. Missing days are not fabricated.

Per-day formula: pull requests closed that day ÷ pull requests opened that day; blank when nothing was opened (a gap, not a zero).

DatePRs openedPRs closedStored score

Releases

Jul 19, 2026

Current0

Previous0

Weekly total

110
Full metrics details

GitHub release published_at events bucketed by UTC day; drafts are excluded.

180 observed daily rows. Missing days are not fabricated.

DateValue

About Lm Evaluation Harness

LM Evaluation Harness from EleutherAI is a framework for few-shot evaluation of language models. Its role in the LLM tooling landscape is straightforward but important: teams need repeatable ways to compare model behavior, and benchmark-style evaluation remains a core part of that process. The project is especially relevant to practitioners who want a more standardized method for measuring performance across tasks or model variants. Unlike prompt workflow tools that focus on application logic, LM Evaluation Harness is centered on assessment. That makes it more useful for model benchmarking, internal comparison, and research-oriented evaluation than for full production orchestration. For teams choosing among models or tracking quality over time, its structured evaluation approach can support more defensible decisions. In short, it is a purpose-built tool for measuring language model capability, not a general LLM app framework.

Join the conversation

Reviews · Questions · Posts

Share what you know about lm-evaluation-harness — write a review from your real experience, ask an implementation question, or publish a post about how you use it.

Share your experience

Write or update your review

Explain what worked, what broke down, and what another team should know before adopting lm-evaluation-harness.

Write your review now — it saves locally and publishes automatically after you sign in.

Rating
Positive attributes
Negative attributes

Project Q&A

Questions and answers

Browse implementation threads tied directly to EleutherAI/lm-evaluation-harness. Each question links through to the full answer page.

Q&A No threads yet

Be the first to ask how teams run lm-evaluation-harness in production. Every question you post becomes a durable, searchable answer page other developers can find.

Ask the first question

Related posts

Posts tagged with the same topics

These posts come from the same topic surface as this repo, so readers can move from project evaluation into practical writeups and migration notes without leaving context.

Posts No topic-linked posts yet

Share how your team uses lm-evaluation-harness — a migration note, an architecture writeup, or a comparison. Your post reaches everyone browsing these same topics.

Write the first post
Get a weekly email with the hottest new projects in the Model Evaluation & Benchmarking and Evaluation world.
No Spam. Unsubscribe easily at any time.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.