§ GREP — First run report

Software mentioned in 1,356 open-access papers, as extracted by the GREP volunteer crowd

What was done between 28 July and 12 August 2026, and every software mention the three models agreed on, with its confidence score.

Report generated 2026-09-04Server data read 2026-09-04 16:07 UTC

Papers in corpus
2,550
Papers complete
1,356
Software mentions
5,665
Distinct names
1,665
Volunteer accounts
8

§ 01 — What we did

GREP (the Great Research Extraction Project) reads scientific papers and records every piece of software they mention: the name, a type (Application, Plugin, ProgrammingEnvironment or OperatingSystem), a purpose (Usage, Creation, Deposition or Mention) and a confidence score from 0 to 1.

Three separately trained models, v1, f13 and f14, read every paper. A mention enters the final record only when at least two of the three find the same span of text. The confidence score on a final mention is the highest score among the models that agreed on it.

The work ran on Lettuce, a volunteer computing system. Volunteers' own machines downloaded packets of about ten papers, ran one of the models offline, and sent back the mentions. Each packet went to up to three different volunteer accounts and was accepted once two of them returned matching output.

  • 2026-07-28Six computations were created on the SciOS Compute server, one per model in a CPU and a GPU version, and a 30-paper test batch was posted to all six.
  • 2026-07-29 → 07-30The test batch completed on CPU machines and on GPU machines. The merged output was identical on both.
  • 2026-07-31The main batch was posted: 2,520 papers in 276 packets. 1,377 papers went to the three CPU computations and 1,143 to the three GPU computations.
  • 2026-07-29 → 08-12Volunteers processed the CPU half. 8 volunteer accounts returned results over the run. The last validation was on 2026-08-12.
  • 2026-08-13One GPU volunteer had processed every f13 and f14 GPU packet once. No second GPU account joined, so those results are not validated.
  • hourly since 07-30A merge job combined the validated outputs of the three models into the final record.

§ 02 — Where the run stands

StagePapers
In the corpus2,550
Read by at least one model and validated1,406
Complete: all three models validated, final record written1,356
Validated by two models, waiting on the third50
PDF could not be read by any model1
On the GPU half: processed once, not yet validated1,143

§ 03 — What we found

5,665 software mentions in the 1,356 complete papers, naming 1,665 distinct software names. 4,240 mentions were found by all three models and 1,425 by two of the three. 576 of the complete papers contain no software mention.

Median confidence 0.786. Before the vote, the models individually found: v1 7,098 mentions in 1,406 papers, f13 5,958 in 1,366, f14 6,262 in 1,396.

By type: Application 4,711, ProgrammingEnvironment 619, Plugin 305, OperatingSystem 30. By purpose: Usage 4,662, Creation 572, Mention 393, Deposition 38.

FieldPapersMentionsDistinct softwarePapers with no mention
Biology2702,36483661
Engineering2361,19439980
Psychology21084622567
Economics20046116090
Environmental science210367134108
Humanities20017075152
Test batch (arts and humanities)302635318
ConfidenceMentions
0.9–1.01,850
0.8–0.9855
0.7–0.8723
0.6–0.7691
0.5–0.6697
0.4–0.5578
0.3–0.4223
0.0–0.348

Software named in the most papers · top 25 of 1,665 names

0255075100125150SPSS130 papersR77 papersExcel47 papersPython36 papersMATLAB26 papersWindows20 papersGraphPad Prism19 papersSmartPLS17 papersGoogle16 papersSPSS Statistics16 papersStata16 papersChatGPT15 papersBLAST15 papersWhatsApp14 papersPROCESS11 papersImageJ11 papersStatistica11 papersggplot211 papersZoom10 papersAMOS9 papersQualtrics9 papersRStudio8 papersscript8 papersBioconductor8 paperslme48 papers

§ 04 — Software list

Every name the models agreed on

Grouped case-insensitively and otherwise as written in the paper. Click a column to sort. The confidence slider hides mentions below the chosen score and recomputes the table.

1,665 names · 5,665 mentions at or above 0.00 · showing 60

SoftwarePapersMentionsMedian confidenceHighestFound by all threeTypePurpose
SPSS1302560.871.0090%ApplicationUsage
R771900.950.9988%ProgrammingEnvironmentUsage
Excel47680.940.9997%ApplicationUsage
Python36790.640.9679%ProgrammingEnvironmentMention
MATLAB261180.800.9995%ProgrammingEnvironmentUsage
Windows20300.650.9743%OperatingSystemUsage
GraphPad Prism19200.960.9885%ProgrammingEnvironmentUsage
SmartPLS17440.951.0093%ApplicationUsage
Google16240.650.9546%ApplicationUsage
SPSS Statistics16210.851.0076%ApplicationUsage
Stata16200.750.9985%ProgrammingEnvironmentUsage
BLAST15340.940.9994%ApplicationUsage
ChatGPT15840.600.9773%ApplicationCreation
WhatsApp14280.690.9775%ApplicationUsage
ggplot211130.770.9769%ApplicationUsage
ImageJ11160.990.9994%ApplicationUsage
PROCESS11270.840.9793%PluginUsage
Statistica11150.820.9787%ApplicationUsage
Zoom10230.880.9978%ApplicationUsage
AMOS9170.960.99100%ApplicationUsage
Qualtrics9110.960.9891%ApplicationUsage
Bioconductor8100.680.9160%ProgrammingEnvironmentUsage
lme48100.940.9780%PluginUsage
RStudio8140.750.8579%ApplicationUsage
script8130.830.9585%PluginUsage
DESeq27130.840.9969%ApplicationUsage
Ensembl7120.790.9867%ApplicationUsage
FastQC7120.960.99100%ApplicationUsage
Google Scholar770.670.910%ApplicationUsage
Google Forms770.930.96100%ApplicationUsage
Mendeley7450.760.9891%ApplicationMention
pandas780.770.9588%PluginUsage
Android6220.850.9673%ApplicationUsage
BioRender690.900.9989%ApplicationUsage
Eviews6140.970.99100%ApplicationUsage
Facebook6140.600.9371%ApplicationUsage
G*Power6100.981.0090%ApplicationUsage
Keras690.640.90100%ApplicationUsage
NVivo6180.900.9956%ApplicationUsage
SAS6120.610.9758%ApplicationUsage
scikit-learn6100.670.9350%ApplicationUsage
Web of Science670.590.9729%ApplicationUsage
Word6120.870.9992%ApplicationUsage
XGBoost6270.590.9574%ApplicationUsage
YouTube6110.620.9355%ApplicationUsage
ClustalW560.860.9783%ApplicationUsage
Cutadapt560.750.98100%ApplicationUsage
featureCounts570.680.97100%ApplicationUsage
JASP580.980.99100%ApplicationUsage
Mplus5130.970.99100%ApplicationUsage
numpy550.830.9160%PluginUsage
Smart PLS550.990.99100%ApplicationUsage
Telegram5330.500.9070%ApplicationUsage
vegan5120.770.9683%ApplicationUsage
Buku440.280.280%ApplicationMention
Dimana460.480.4983%ApplicationMention
edgeR4130.710.9677%ApplicationUsage
emmeans480.760.9688%PluginUsage
Ethereum4150.580.8940%ApplicationUsage
figshare440.760.9075%ApplicationUsage

§ 05 — By paper

The 1,356 papers with a complete record

Each chip is one mention with its confidence. Three dots mean all three models found it; two dots mean two of three.

Loading the per-paper data…

●●● all three models●●○ two of threenumber = confidence

§ 06 — Technical appendix

The six computations (leaves) on infra.scios.tech
LeafHardwarePacketsValidatedQueuedFailedResultsAccountsFirst resultLast validation
extract2-student-crowd-v1CPU1421420028562026-07-292026-08-12
extract2-student-crowd-v1-gpuGPU140313703032026-07-302026-07-30
extract2-student-crowd-f13CPU1421381327852026-07-292026-08-12
extract2-student-crowd-f13-gpuGPU1403137014332026-07-302026-07-30
extract2-student-crowd-f14CPU1421411028462026-07-292026-08-12
extract2-student-crowd-f14-gpuGPU1403137014332026-07-302026-07-30
Totals
846 packets · 1,163 results · 8 volunteer accounts
Images
ghcr.io/jring-o/extract2-student:2.1-{v1,f13,f14} and -gpu
Validation
3 target copies, quorum 2, numeric tolerance 0.01, timing fields ignored, up to 3 retries
Packets
up to 10 papers and under 18 MB each, fetched by volunteers from a public bucket by SHA-256 name
Models
DeBERTa-v3-large students distilled from the Extract2 teacher ensemble; same weights on CPU and GPU
Agreement between the three models
Papers compared
1,356
Spans found by all three
4,240
By two models
1,425
By one model only (dropped by the vote)
3,489
Pairwise span overlap (Jaccard)
v1-f13 0.5343 · v1-f14 0.531 · f13-f14 0.7292
Same type on shared spans
96%
Same purpose on shared spans
85%
Not in the final record
  • 50 papers are validated by two models and wait on the third: 40 on f13 (three failed packets and one queued packet), 10 on f14 (one queued packet).
  • 1 paper (259676743) could not be read by any model: invalid PDF.
  • 1,143 papers on the GPU half have one unvalidated result each from f13 and f14 (and 210 from v1). They are excluded here.
  • 3 results were rejected in validation, all on one paper (251609378), where the paragraph count differed between machines; the paper's mentions were unaffected.
Download the tables
  • software_by_package.csv — every name: papers, mentions, median and highest confidence, share found by all three models, type, purpose
  • software_by_paper.csv — every mention with its paper, score and vote count
  • papers.csv — all 2,550 papers in the corpus with title, field and status
GREP first run · SciOS Compute