Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Repository files navigation
REsearch POOLer (repool) Project site: https://sites.google.com/site/researchpooler/home Web interface (Stage 4 - live since 2026): https://conftrace.com Authors: Andrej Karpathy <karpathy@cs.stanford.edu> || <andrej.karpathy@gmail.com>, http://cs.stanford.edu/~karpathy/ Stage 4 web UI by: Justyna Wojtczak <justine84@gmail.com> (github.com/justi) ------------------------------------------------------------------------------- MOTIVATION AND PLAN ------------------------------------------------------------------------------- - Ever wish you could right away view all papers published based on a keyword in title or abstract? - Or, ever wish you could look up the most similar paper (content wise) to some paper on some random url? - How about searching for all papers that report a result on a particular dataset? -> Literature review is much harder than it should be. This set of tools is an initiative to fix this problem. Here's the master plan and types of scripts in this project: STAGE 1 scripts: scripts for raw data gathering and parsing that output pickles in an intermediate dictionary-based representation. These will include mostly scripts that download files, parse HTML, etc. STAGE 2 scripts: scripts that enrich the intermediate representations from STAGE 1 in various ways. For example, a script could iterate over publications in database, and if it finds that some entry is missing its pdf contents, it could attempt a google search to find the pdf, and add it if successful. STAGE 3 scripts: tools and helper functions that can analyze the intermediate representations and produce higher level scripts that do more interesting things. For example, functionality such as 'find all documents that are similar to this one', or 'find object detection papers in psychology'. All kinds of fun Machine Learning can go here as well, like LDA etc. STAGE 4 - DONE (2026): web-based UI at https://conftrace.com - Browse 220k papers from 36 top CS/AI conferences (NeurIPS, ICML, ICLR, CVPR, ACL, ...) - LLM-classified taxonomy of 290+ topics - 17k keyword cloud with co-occurrence graph (Neo4j-powered /explore) - Per-author profiles + 15 achievement badges - Emerging terms detector + conference DNA - Sitemap with 630k URLs for SEO - Source: github.com/justi/research-explorer (Rails 8 + PostgreSQL + Neo4j) ------------------------------------------------------------------------------- INSTALLATION ------------------------------------------------------------------------------- pip install -r requirements.txt Then run any scraper to generate its database: $> python nips_download_parse.py # scrape metadata $> python nips_add_pdftext.py # extract PDF text (slow) $> python nips_add_abstracts.py # NeurIPS abstracts (legacy, separate) $> python add_abstracts.py --status # multi-source abstract scraper $> python add_abstracts.py --source acl_anthology # 7 source plugins: acl_anthology, pmlr, cvf, openreview, isca, jmlr, usenix, aaai, rss ------------------------------------------------------------------------------- EXAMPLE USAGE ------------------------------------------------------------------------------- Say you unexpectedly became very interested in Deep Belief Networks. It now takes 3 lines of python to open all NIPS papers that mention 'deep' in title inside your browser: (also see demo2) >>> pubs = loadPubs('pubs_nips') >>> p = [x['pdf'] for x in pubs if 'deep' in x['title'].lower()] >>> openPDFs(p) Or maybe you want to open all papers that mention MNIST dataset? (demo1 also shows how you can easily go on to open the 3 latest ones.) >>> pubs = loadPubs('pubs_nips') >>> p = [x['title'] for x in pubs if 'mnist' in x.get('pdf_text',{})] >>> openPDFs(p) Or how about opening papers that are most similar to some paper at some url? See demo3. ------------------------------------------------------------------------------- AVAILABLE CONFERENCES (36 scrapers, 220,881 papers) ------------------------------------------------------------------------------- Scraper Conference Source Years Papers ---------------------------- ----------- ------------------------- --------- ------ nips_download_parse.py NeurIPS proceedings.neurips.cc 2006-2024 21,859 acl_download_parse.py ACL aclanthology.org 2000-2025 21,480 emnlp_download_parse.py EMNLP aclanthology.org 2000-2025 20,631 aaai_download_parse.py AAAI ojs.aaai.org 2019-2026 20,013 cvpr_download_parse.py CVPR openaccess.thecvf.com 2013-2025 18,452 icml_download_parse.py ICML proceedings.mlr.press 2013-2025 14,281 iclr_download_parse.py ICLR openreview.net 2018-2025 11,015 ijcai_download_parse.py IJCAI ijcai.org 2013-2025 9,941 iccv_download_parse.py ICCV openaccess.thecvf.com 2013-2025 9,145 naacl_download_parse.py NAACL aclanthology.org 2000-2025 8,918 interspeech_download_parse.py INTERSPEECH isca-archive.org 2016-2024 8,778 coling_download_parse.py COLING aclanthology.org 2000-2025 8,740 eccv_download_parse.py ECCV ecva.net 2018-2024 6,166 ijcnlp_download_parse.py IJCNLP aclanthology.org 2005-2025 5,328 aistats_download_parse.py AISTATS proceedings.mlr.press 2010-2025 4,609 eacl_download_parse.py EACL aclanthology.org 2003-2026 4,564 wacv_download_parse.py WACV openaccess.thecvf.com 2020-2026 4,435 jmlr_download_parse.py JMLR jmlr.org 2000-2025 4,133 semeval_download_parse.py SemEval aclanthology.org 2007-2025 3,217 miccai_download_parse.py MICCAI papers.miccai.org 2024-2025 1,883 colt_download_parse.py COLT proceedings.mlr.press 2011-2025 1,594 conll_download_parse.py CoNLL aclanthology.org 2000-2025 1,547 corl_download_parse.py CoRL proceedings.mlr.press 2017-2025 1,488 rss_download_parse.py RSS roboticsproceedings.org 2005-2025 1,469 uai_download_parse.py UAI proceedings.mlr.press 2019-2025 1,368 aacl_download_parse.py AACL aclanthology.org 2020-2025 1,333 acml_download_parse.py ACML proceedings.mlr.press 2010-2024 825 nsdi_download_parse.py NSDI usenix.org 2012-2025 821 l4dc_download_parse.py L4DC proceedings.mlr.press 2020-2025 666 midl_download_parse.py MIDL proceedings.mlr.press 2019-2024 499 osdi_download_parse.py OSDI usenix.org 2012-2025 472 alt_download_parse.py ALT proceedings.mlr.press 2017-2025 381 mlhc_download_parse.py MLHC proceedings.mlr.press 2016-2025 323 pgm_download_parse.py PGM proceedings.mlr.press 2016-2024 220 clear_download_parse.py CLeaR proceedings.mlr.press 2022-2025 190 automl_download_parse.py AutoML proceedings.mlr.press 2016-2025 97 TOTAL 220,881 ------------------------------------------------------------------------------- TAXONOMY SYSTEM (taxonomy/) ------------------------------------------------------------------------------- LLM-powered topic classification. Papers are tagged with topics from a controlled 3-level hierarchy (209 categories in taxonomy/config.yaml) and normalized keywords. Features: - Chain-of-Thought reasoning (inspired by TopicGPT) — the LLM explains its classification choices by citing words from the title/abstract - Abstract-aware — when abstracts are available (e.g. via nips_add_abstracts.py), they are included in the prompt for more accurate classification - New topic suggestion — if no existing path fits, the LLM can propose a new topic path prefixed with "Other:" instead of forcing a bad match Setup: python nips_add_abstracts.py # (optional) scrape NeurIPS abstracts python taxonomy/db.py --init # create SQLite DB with all papers python taxonomy/classify.py --conf nips # classify papers (uses OpenCode CLI) python taxonomy/db.py --import-taxonomy # import classifications into DB Options: python taxonomy/classify.py --resume # resume from checkpoint python taxonomy/classify.py --abstract-only # only papers with abstracts python taxonomy/classify.py --build-index # build topic/keyword indexes Browse: python taxonomy/db.py --stats # show statistics python taxonomy/db.py --tree --depth 2 # topic hierarchy tree python taxonomy/db.py --topic "Deep Learning" # papers by topic python taxonomy/db.py --keyword "attention" # papers by keyword python taxonomy/db.py --related "transformers" # co-occurring keywords python taxonomy/db.py --query "transformer" # search by title ------------------------------------------------------------------------------- ORGANIZATION, I/O AND DATA REPRESENTATIONS ------------------------------------------------------------------------------- Here's the idea for the near future, I think: there will be several stage1 scripts, each of which is reponsible for parsing a particular venue of publications. For example, the stage1 script nips_download_parse.py parses and outputs all publications in NIPS from 2003 to 2011. (but does not analyze the text) The idea is to have very similar scripts for other venues, such as ICML, or CVPR, etc... The output of each such script should be a pickled list of dictionaries. Each dictionary represents a publication. For example: [{'title': 'Solving AI using Random Forests', 'authors': ['Jim Smith', 'Bill Smith', 'year': 2020, 'venue': 'NIPS 2020', 'pdf': 'http://google.com/ai'}, ...] this representation is a flexible start, as some conference pages provide more information than others, and we don't want to force any particular structure from the get go. In other words, the database could contain some papers that have the author, title, and abstract, but not the full text. Another entry might have the full text, but maybe it is missing author or title. Stage 2 scripts will be useful to go over these representations, and fill in details in whatever ways possible. (such as maybe hooking into other sites like Google Scholar, etc?) Note: I am well aware that the "flat list of dictionaries pickled in a file" representation isn't scalable. However, I am a believer of avoiding premature encapsulation. Goal is to keep things as flat as possible, as long as possible, and to avoid immediate over-engineering of things. Lastly, this representation is actually kinda neat because it lets you run all kinds of nice queries very quickly using list comprehensions. For example: #all papers by Andrew Ng >>> [x['title'] for x in pubs if any("A. Ng" in a for a in x['authors'])] ------------------------------------------------------------------------------- PIPELINE ------------------------------------------------------------------------------- STAGE 1: Scrapers (*_download_parse.py) fetch proceedings and save as pickles. STAGE 2: Enrichment scripts add content to papers: - nips_add_pdftext.py: downloads PDFs and extracts bag-of-words text - nips_add_abstracts.py: extracts abstracts from NeurIPS HTML pages (no PDF download needed — much faster than pdf_text extraction) STAGE 3: repool_analysis.py provides similarity search across publications. STAGE 4: taxonomy/ classifies papers into hierarchical topics + keywords. ------------------------------------------------------------------------------- FAQ ------------------------------------------------------------------------------- Q: Your file x does this, but that's bad practice, and you want to do y. Also you have a bug in z. A: Most likely agreed. These scripts/functions are something I hacked together in 3 days during periods of about 1am to 6am. I'd be happy to hear thoughs/suggestions or see fixes or better/alternative ways of doing things. Check the website for this project, which comes with discussion forum attached. ------------------------------------------------------------------------------- MANPOWER ADVERTISEMENTS / IDEAS ------------------------------------------------------------------------------- Advertisement posting 1: 1. Pick your favorite conference/journal 2. Look through their page and write an HTML parser in style of my nips_download_parse.py 3. Output the same type of representation as described above in a file 'pubs_conferenceblah' 4. Publish the scripts you used so that it is faster for others to do similar things 5. Upload your output pubs_ file, or send it to me for publishing Advertisement posting 2: It should be possible to take arbitrary directory full of PDF files, and create pubs_ file for them. Can Title, Authors be reliably extracted from PDFs somehow? Can tools be made that at least partially automate the process so that different parsers don't have to be written for each venue? Can we find large databases of papers/information online that we can scrape and enter? Advertisement posting 3: Heavy duty Machine Learning tools (such as Naive Bayes, lol) needed that can work on top of representations stored in pubs_ files and answer questions such as: 'what papers in database are most similar to the one on this url?', or, 'What are the common topics?' etc etc... ------------------------------------------------------------------------------- DEPENDENCIES ------------------------------------------------------------------------------- beautifulsoup4 HTML parsing pdfminer.six PDF text extraction (for stage 2) pyyaml Taxonomy config parsing (for taxonomy/)