Skip to content
 
 

Latest commit

 

History

72 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

REsearch POOLer (repool)
Project site: https://sites.google.com/site/researchpooler/home
Web interface (Stage 4 - live since 2026): https://conftrace.com

Authors: Andrej Karpathy <karpathy@cs.stanford.edu> || <andrej.karpathy@gmail.com>, http://cs.stanford.edu/~karpathy/
Stage 4 web UI by: Justyna Wojtczak <justine84@gmail.com> (github.com/justi)

-------------------------------------------------------------------------------
MOTIVATION AND PLAN
-------------------------------------------------------------------------------
- Ever wish you could right away view all papers published based on a keyword in title or abstract?
- Or, ever wish you could look up the most similar paper (content wise) to some paper on some random url?
- How about searching for all papers that report a result on a particular dataset?
-> Literature review is much harder than it should be.

This set of tools is an initiative to fix this problem. Here's the master plan and types of scripts in this project:

STAGE 1 scripts: scripts for raw data gathering and parsing that output pickles in an intermediate dictionary-based representation. These will include mostly scripts that download files, parse HTML, etc.
STAGE 2 scripts: scripts that enrich the intermediate representations from STAGE 1 in various ways. For example, a script could iterate over publications in database, and if it finds that some entry is missing its pdf contents, it could attempt a google search to find the pdf, and add it if successful.
STAGE 3 scripts: tools and helper functions that can analyze the intermediate representations and produce higher level scripts that do more interesting things. For example, functionality such as 'find all documents that are similar to this one', or 'find object detection papers in psychology'. All kinds of fun Machine Learning can go here as well, like LDA etc.
STAGE 4 - DONE (2026): web-based UI at https://conftrace.com
  - Browse 220k papers from 36 top CS/AI conferences (NeurIPS, ICML, ICLR, CVPR, ACL, ...)
  - LLM-classified taxonomy of 290+ topics
  - 17k keyword cloud with co-occurrence graph (Neo4j-powered /explore)
  - Per-author profiles + 15 achievement badges
  - Emerging terms detector + conference DNA
  - Sitemap with 630k URLs for SEO
  - Source: github.com/justi/research-explorer (Rails 8 + PostgreSQL + Neo4j)

-------------------------------------------------------------------------------
INSTALLATION
-------------------------------------------------------------------------------
pip install -r requirements.txt

Then run any scraper to generate its database:
$> python nips_download_parse.py                         # scrape metadata
$> python nips_add_pdftext.py                            # extract PDF text (slow)
$> python nips_add_abstracts.py                          # NeurIPS abstracts (legacy, separate)
$> python add_abstracts.py --status                      # multi-source abstract scraper
$> python add_abstracts.py --source acl_anthology        # 7 source plugins: acl_anthology, pmlr, cvf, openreview, isca, jmlr, usenix, aaai, rss

-------------------------------------------------------------------------------
EXAMPLE USAGE
-------------------------------------------------------------------------------
Say you unexpectedly became very interested in Deep Belief Networks. It now takes 3 lines of python to open all NIPS papers that mention 'deep' in title inside your browser: (also see demo2)

>>> pubs = loadPubs('pubs_nips')
>>> p = [x['pdf'] for x in pubs if 'deep' in x['title'].lower()]
>>> openPDFs(p)

Or maybe you want to open all papers that mention MNIST dataset? (demo1 also shows how you can easily go on to open the 3 latest ones.)
>>> pubs = loadPubs('pubs_nips')
>>> p = [x['title'] for x in pubs if 'mnist' in x.get('pdf_text',{})]
>>> openPDFs(p)

Or how about opening papers that are most similar to some paper at some url? See demo3.

-------------------------------------------------------------------------------
AVAILABLE CONFERENCES (36 scrapers, 220,881 papers)
-------------------------------------------------------------------------------
  Scraper                        Conference    Source                      Years       Papers
  ----------------------------   -----------   -------------------------   ---------   ------
  nips_download_parse.py         NeurIPS       proceedings.neurips.cc      2006-2024   21,859
  acl_download_parse.py          ACL           aclanthology.org            2000-2025   21,480
  emnlp_download_parse.py        EMNLP         aclanthology.org            2000-2025   20,631
  aaai_download_parse.py         AAAI          ojs.aaai.org                2019-2026   20,013
  cvpr_download_parse.py         CVPR          openaccess.thecvf.com       2013-2025   18,452
  icml_download_parse.py         ICML          proceedings.mlr.press       2013-2025   14,281
  iclr_download_parse.py         ICLR          openreview.net              2018-2025   11,015
  ijcai_download_parse.py        IJCAI         ijcai.org                   2013-2025    9,941
  iccv_download_parse.py         ICCV          openaccess.thecvf.com       2013-2025    9,145
  naacl_download_parse.py        NAACL         aclanthology.org            2000-2025    8,918
  interspeech_download_parse.py  INTERSPEECH   isca-archive.org            2016-2024    8,778
  coling_download_parse.py       COLING        aclanthology.org            2000-2025    8,740
  eccv_download_parse.py         ECCV          ecva.net                    2018-2024    6,166
  ijcnlp_download_parse.py       IJCNLP        aclanthology.org            2005-2025    5,328
  aistats_download_parse.py      AISTATS       proceedings.mlr.press       2010-2025    4,609
  eacl_download_parse.py         EACL          aclanthology.org            2003-2026    4,564
  wacv_download_parse.py         WACV          openaccess.thecvf.com       2020-2026    4,435
  jmlr_download_parse.py         JMLR          jmlr.org                    2000-2025    4,133
  semeval_download_parse.py      SemEval       aclanthology.org            2007-2025    3,217
  miccai_download_parse.py       MICCAI        papers.miccai.org           2024-2025    1,883
  colt_download_parse.py         COLT          proceedings.mlr.press       2011-2025    1,594
  conll_download_parse.py        CoNLL         aclanthology.org            2000-2025    1,547
  corl_download_parse.py         CoRL          proceedings.mlr.press       2017-2025    1,488
  rss_download_parse.py          RSS           roboticsproceedings.org     2005-2025    1,469
  uai_download_parse.py          UAI           proceedings.mlr.press       2019-2025    1,368
  aacl_download_parse.py         AACL          aclanthology.org            2020-2025    1,333
  acml_download_parse.py         ACML          proceedings.mlr.press       2010-2024      825
  nsdi_download_parse.py         NSDI          usenix.org                  2012-2025      821
  l4dc_download_parse.py         L4DC          proceedings.mlr.press       2020-2025      666
  midl_download_parse.py         MIDL          proceedings.mlr.press       2019-2024      499
  osdi_download_parse.py         OSDI          usenix.org                  2012-2025      472
  alt_download_parse.py          ALT           proceedings.mlr.press       2017-2025      381
  mlhc_download_parse.py         MLHC          proceedings.mlr.press       2016-2025      323
  pgm_download_parse.py          PGM           proceedings.mlr.press       2016-2024      220
  clear_download_parse.py        CLeaR         proceedings.mlr.press       2022-2025      190
  automl_download_parse.py       AutoML        proceedings.mlr.press       2016-2025       97
                                                                           TOTAL      220,881

-------------------------------------------------------------------------------
TAXONOMY SYSTEM (taxonomy/)
-------------------------------------------------------------------------------
LLM-powered topic classification. Papers are tagged with topics from a
controlled 3-level hierarchy (209 categories in taxonomy/config.yaml)
and normalized keywords.

Features:
 - Chain-of-Thought reasoning (inspired by TopicGPT) — the LLM explains its
   classification choices by citing words from the title/abstract
 - Abstract-aware — when abstracts are available (e.g. via nips_add_abstracts.py),
   they are included in the prompt for more accurate classification
 - New topic suggestion — if no existing path fits, the LLM can propose a new
   topic path prefixed with "Other:" instead of forcing a bad match

Setup:
   python nips_add_abstracts.py              # (optional) scrape NeurIPS abstracts
   python taxonomy/db.py --init              # create SQLite DB with all papers
   python taxonomy/classify.py --conf nips   # classify papers (uses OpenCode CLI)
   python taxonomy/db.py --import-taxonomy   # import classifications into DB

Options:
   python taxonomy/classify.py --resume           # resume from checkpoint
   python taxonomy/classify.py --abstract-only    # only papers with abstracts
   python taxonomy/classify.py --build-index      # build topic/keyword indexes

Browse:
   python taxonomy/db.py --stats             # show statistics
   python taxonomy/db.py --tree --depth 2    # topic hierarchy tree
   python taxonomy/db.py --topic "Deep Learning"    # papers by topic
   python taxonomy/db.py --keyword "attention"      # papers by keyword
   python taxonomy/db.py --related "transformers"   # co-occurring keywords
   python taxonomy/db.py --query "transformer"      # search by title

-------------------------------------------------------------------------------
ORGANIZATION, I/O AND DATA REPRESENTATIONS
-------------------------------------------------------------------------------

Here's the idea for the near future, I think: there will be several stage1 scripts, each of which is reponsible for parsing a particular venue of publications. For example, the stage1 script nips_download_parse.py parses and outputs all publications in NIPS from 2003 to 2011. (but does not analyze the text)

The idea is to have very similar scripts for other venues, such as ICML, or CVPR, etc... The output of each such script should be a pickled list of dictionaries. Each dictionary represents a publication. For example:
[{'title': 'Solving AI using Random Forests', 'authors': ['Jim Smith', 'Bill Smith', 'year': 2020, 'venue': 'NIPS 2020', 'pdf': 'http://google.com/ai'},
...]

this representation is a flexible start, as some conference pages provide more information than others, and we don't want to force any particular structure from the get go. In other words, the database could contain some papers that have the author, title, and abstract, but not the full text. Another entry might have the full text, but maybe it is missing author or title. Stage 2 scripts will be useful to go over these representations, and fill in details in whatever ways possible. (such as maybe hooking into other sites like Google Scholar, etc?)

Note: I am well aware that the "flat list of dictionaries pickled in a file" representation isn't scalable. However, I am a believer of avoiding premature encapsulation. Goal is to keep things as flat as possible, as long as possible, and to avoid immediate over-engineering of things.

Lastly, this representation is actually kinda neat because it lets you run all kinds of nice queries very quickly using list comprehensions. For example:

#all papers by Andrew Ng
>>> [x['title'] for x in pubs if any("A. Ng" in a for a in x['authors'])]

-------------------------------------------------------------------------------
PIPELINE
-------------------------------------------------------------------------------
STAGE 1: Scrapers (*_download_parse.py) fetch proceedings and save as pickles.
STAGE 2: Enrichment scripts add content to papers:
         - nips_add_pdftext.py: downloads PDFs and extracts bag-of-words text
         - nips_add_abstracts.py: extracts abstracts from NeurIPS HTML pages
           (no PDF download needed — much faster than pdf_text extraction)
STAGE 3: repool_analysis.py provides similarity search across publications.
STAGE 4: taxonomy/ classifies papers into hierarchical topics + keywords.

-------------------------------------------------------------------------------
FAQ
-------------------------------------------------------------------------------
Q: Your file x does this, but that's bad practice, and you want to do y. Also you have a bug in z.
A: Most likely agreed. These scripts/functions are something I hacked together in 3 days during periods of about 1am to 6am. I'd be happy to hear thoughs/suggestions or see fixes or better/alternative ways of doing things. Check the website for this project, which comes with discussion forum attached.

-------------------------------------------------------------------------------
MANPOWER ADVERTISEMENTS / IDEAS
-------------------------------------------------------------------------------
Advertisement posting 1:
1. Pick your favorite conference/journal
2. Look through their page and write an HTML parser in style of my nips_download_parse.py
3. Output the same type of representation as described above in  a file 'pubs_conferenceblah'
4. Publish the scripts you used so that it is faster for others to do similar things
5. Upload your output pubs_ file, or send it to me for publishing

Advertisement posting 2:
It should be possible to take arbitrary directory full of PDF files, and create pubs_ file for them. Can Title, Authors be reliably extracted from PDFs somehow? Can tools be made that at least partially automate the process so that different parsers don't have to be written for each venue? Can we find large databases of papers/information online that we can scrape and enter?

Advertisement posting 3:
Heavy duty Machine Learning tools (such as Naive Bayes, lol) needed that can work on top of representations stored in pubs_ files and answer questions such as: 'what papers in database are most similar to the one on this url?', or, 'What are the common topics?'

etc etc...

-------------------------------------------------------------------------------
DEPENDENCIES
-------------------------------------------------------------------------------
beautifulsoup4    HTML parsing
pdfminer.six      PDF text extraction (for stage 2)
pyyaml            Taxonomy config parsing (for taxonomy/)

About

Research paper discovery & taxonomy system. 214k papers from 36 AI/ML/NLP/CV conferences with LLM-powered topic classification.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages