Projects · Case study · 01
Contributable
Find an open source project that answers newcomers. Contributable indexes the repositories of every Google Summer of Code organisation from 2024 to 2026 and measures how each treats people outside its core team: how fast a person replies to a first pull request, and how often it gets merged. A scheduled job keeps the index fresh; every figure shows its sample and links to the pull requests behind it. No blended score, no AI, no cost to run.
Open the app ↗GSoC ranking ↗Repository ↗
Architecture
The system, at a glance.
01 — Claims and evidence
| Claim | Measurement | Context | Verify |
|---|---|---|---|
| Index size | 1,575 repositories | 253 GSoC organisations from 2024 to 2026, 111 languages; live status on 2026-10-04 | live status API ↗ |
| Freshness | sweep 4× an hour | every indexed repository was refreshed within two days on 2026-10-04; each run spends one hour’s allowance, stalest first | status page ↗ |
| Test suite | 210 tests · 12 files | survival estimate, merge and reply metrics, starter issue states, trends, the explore query, the GSoC ranking rule, matching and discovery; passing from a fresh clone on 2026-10-04 | README · testing ↗ |
| Reply-time estimate | Kaplan-Meier | a pull request nobody answered stays in the estimate as still waiting, so ignoring people makes the figure worse instead of dropping out | methodology ↗ |
| Sample-size rule | under 5 → “Not enough data” | every figure shows its sample and links to the pull requests or issues it counted | README · what is measured ↗ |
| Open data and API | CC BY 4.0 · /api/v1 | the full index on a data branch, a read-only API, README badges and Atom feeds of available issues | data branch ↗ |
| Running cost | none | no database, no paid API, no card on file: GitHub Actions for the job, static files and cached reads on Vercel’s free tier | README · cost ↗ |
| Scale of the build | 136 commits · 15 pages | TypeScript, Next.js 16, React 19, three.js; code under AGPL-3.0 with an author-credit notice | commit history ↗ |
02 — The problem
New contributors pick the most famous project, open a pull request, and wait. Research on newcomer onboarding keeps finding the same two obstacles: nobody replies, and it is hard to find a task that is not already taken. Stars say nothing about either.
Contributable measures both and lets you search by them. It indexes the repositories of every Google Summer of Code organisation from 2024 to 2026, measures how each one treats people outside its core team, and publishes the result as a website, an open dataset and an API. Any other public repository can be checked on the spot.
It started as a single-repository report and grew in four days into an index: a scheduled pipeline, fifteen pages, a public API, README badges and Atom feeds.
03 — How it works
From the GSoC programme list to a page that never waits on GitHub
- The universe: the organisations of every GSoC year from 2024 to 2026, mapped to their GitHub accounts. Umbrella organisations whose work lives elsewhere are mapped by hand, and those that cannot be mapped are listed as not measured rather than ranked last.
- Discovery picks up to 12 of each organisation’s most starred repositories that are not forks, archives or mirrors, were pushed in the last 180 days and have at least 5 pull requests.
- A scheduled job in GitHub Actions starts four times an hour. It reads the GitHub GraphQL API with the token Actions provides, spends that hour’s allowance on the stalest repositories, and stops before the limit. A first read fetches a year of pull requests and issues; later reads fetch only what changed.
- Pure functions turn that into figures for pull requests from outside contributors: the merge rate, the first reply time estimated with Kaplan-Meier so unanswered pull requests count as still waiting, time to merge, starter issue states, reply hours and a four-week trend. Bots are removed by account type, name and behaviour.
- The job commits everything to a data branch as one snapshot, published under CC BY 4.0. Author names are stored only as one-way hashes.
- The Next.js site on Vercel reads those files through the framework’s fetch cache, so a page view never calls GitHub and never waits on it. The same data serves Explore, the GSoC ranking, repository pages, the issue finder, matching, comparison, the API, README badges and Atom feeds.
No blended score. Every ranking states its rule on the page.
04 — Decisions
A scheduled pipeline, not live requests
The first version read GitHub when someone opened a report, which tied every page view to a rate limit. Now a job in GitHub Actions does all the reading and commits the result to a data branch; the website only reads files. Pages are fast, the free API allowance is spent on keeping the index fresh, and the whole system costs nothing to run.
Count the people still waiting
A plain median of reply times leaves out pull requests nobody answered, so a project that ignores half its contributors can look quick. The Kaplan-Meier estimate, the standard way to measure a wait that has not ended, keeps them in as still waiting. If more than half were never answered there is no median at all.
No blended score, and every ranking states its rule
The GSoC ranking orders organisations by the share of pull requests answered within 7 days, then by outside merge rate, and prints that rule on the page. Nothing is rolled into one health number, because there is no defensible way to weigh the signals for every project.
Describe projects, never people
Author names are stored as one-way hashes and nothing is published per person. Reply hours pool the whole team and are withheld when fewer than three people replied, because the chart would then be one maintainer’s schedule. Maintainers can ask for a repository to be removed.
A cohort old enough to judge
Merge rates use pull requests opened 30 to 120 days ago. Newer ones have not had time to be answered or merged, and including them would make every project look worse than it is.
Open by default
The dataset is on a public branch under CC BY 4.0 with a read-only API, so anyone can check a figure against GitHub or build on it. The code is AGPL-3.0, so a copy run as a website must publish its source and keep the author credit.
05 — What went wrong
Stated on purpose.
Bots made projects look fast
Bots often post the first comment on a pull request, and plenty run on ordinary accounts that GitHub does not mark as bots. Detection now works on three levels: account type, name (CI, CLA and -bot accounts), and behaviour, so an account that answers within a minute on most pull requests, or posts the same templated message again and again, no longer counts as a reply.
GSoC organisations are not GitHub organisations
Some GSoC organisations are umbrellas whose projects live under several other GitHub accounts, and some do not work on GitHub at all. They now go through a hand-checked mapping file; the ones that cannot be mapped are listed as not measured instead of being ranked last for silence.
Small organisations topped the ranking
An organisation with a handful of pull requests could post a perfect reply rate by chance. A place in the ranking now needs 20 outside pull requests; smaller organisations are listed after it with their figures.
Changed definitions left stale results
When a definition changed, repositories already stored were still shown under the old rules until their turn came round. Results now record which rules produced them, are re-read when those rules change, and a share of each run’s budget is reserved for those full re-reads.
06 — Limitations
- Projects that review on a mailing list, Gerrit or GitLab look silent here; organisations that do not work on GitHub are listed as not measured.
- A bot that fits none of the detection rules still counts as a person and makes a project look faster than it is.
- GitHub only marks someone as a member when their membership is public, so a core developer with private membership counts as an outside contributor.
- A reply from another outside contributor counts as a human reply. It is an answer, but not a maintainer’s.
- Claims on starter issues are detected from English comment wording; a claim written differently is missed.
- A pull request answered only by an approving review, without a comment, looks unanswered.
- Tone is not measured. Nothing here tells you whether a community is welcoming.
07 — What I’d do next
- Read approving reviews, so pull requests answered only by a review count as answered.
- Measure projects that review on GitLab or Gerrit instead of listing them as not measured.
- Detect claims on starter issues written in languages other than English.
Last verified 2026-10-04 · numbers reported as measured, with their context and source
