CiteStart

Nobody tells you what to read first.

You start a PhD with a topic and forty years of literature behind it. Describe your topic and get a short list of the papers that actually matter, in the order to read them, with a reason for each one.

Free preview first. The full report is €15, paid once, no account needed.

The topic that produced it

“Deep learning methods for predicting protein structure from amino acid sequence, including AlphaFold and related approaches”

That sentence is the whole input. Everything beside it was produced from it, in under four minutes.

Research starter report

Deep Learning Approaches for Protein Structure Prediction from Sequence

Field
Biology / Life Sciences
Papers
15, every DOI checked against Crossref
Sources
Semantic Scholar, OpenAlex, Crossref

Field overview

Determining the three-dimensional structure of a protein from its one-dimensional amino acid sequence, known as the protein folding problem, has been one of the central challenges in molecular biology for over half a century.

Reading order

  1. Before and after AlphaFold2: An overview of protein structure prediction

    Letícia Machado Favery Bertoline et al., 2023. Frontiers in Bioinformatics. doi:10.3389/fbinf.2023.1120370 312 citations

    Start here: this historical overview of protein structure prediction before and after AlphaFold2 gives essential context for understanding where the field came from and why the deep learning revolution mattered.

  2. Protein structure prediction via deep learning: an in-depth review

    Yajie Meng et al., 2025. Frontiers in Pharmacology. doi:10.3389/fphar.2025.1498662 39 citations

    Read after the historical overview, as this in-depth review of deep learning methods for protein structure prediction consolidates the broader technical landscape and prepares you for diving into specific architectures.

An excerpt from a real report. Read all 15 papers.

How it works

  1. 1

    You describe your topic

    A few sentences in plain language. Mention the methods or problems you care about, and any papers or authors you already know, and those get used as starting points.

  2. 2

    We search and check

    Your topic becomes a set of search queries against Semantic Scholar, OpenAlex and Crossref. Candidates are scored for how directly they address your topic, and every DOI is looked up in Crossref to confirm it resolves to a real record.

  3. 3

    You get a reading list

    Up to fifteen papers in a deliberate order, each with a short summary, a reason to read it, and where it sits in the field. Plus trends, gaps, open questions and next steps. Download it as a PDF.

Generating a report takes two to four minutes.

In the report

The papers are the smallest part of it.

Every paper carries a reason to read it at that position, and a summary written against the paper's title rather than a scraped abstract.

Reading order, no. 1 of 15

Before and after AlphaFold2: An overview of protein structure prediction

Letícia Machado Favery Bertoline et al., 2023. Frontiers in Bioinformatics.

Start here: this historical overview of protein structure prediction before and after AlphaFold2 gives essential context for understanding where the field came from and why the deep learning revolution mattered.

Where the literature is thin, so you can aim your own work at something unanswered rather than something already settled.

Research gaps

Reliable Prediction of Intrinsically Disordered Regions and Proteins

Current top-performing models such as AlphaFold (Jumper et al., 2021) were primarily benchmarked on well-folded globular proteins, and their confidence metrics do not straightforwardly translate to intrinsically disordered proteins or regions. Despite acknowledgment of this limitation in reviews by Meng et al. (2025) and Szelogowski et al. (2025), systematic methods for predicting the ensemble behavior of disordered regions remain underexplored.

Questions you could actually take to a supervisor meeting, drawn from the papers in your own list.

Open questions

Can a deep learning model be trained or fine-tuned to predict multiple distinct conformational states of a protein from sequence alone, and how should training data be curated to capture functionally relevant structural diversity rather than crystallographic noise?

And what to do on Monday morning. Methods and datasets, keywords for your own searching, and concrete next steps.

Next steps

Reproduce a small-scale AlphaFold inference run using ColabFold (Mirdita et al., 2022), which provides free GPU access via Google Colab. Pick a well-characterized protein with a known crystal structure from the PDB, predict its structure, and quantitatively compare your prediction to the experimental structure using TM-score and RMSD metrics to build intuition for what 'good' predictions look like in practice.

Why trust it

Nothing is invented.

The papers are real
The papers and their details come straight from Semantic Scholar, OpenAlex and Crossref. Nothing in the list is recalled from a language model's memory. The summaries, reading order and explanations are written by AI from each paper's abstract, or from its title where no reliable abstract exists, so read them as a guide, not as the paper.
Every DOI is checked
Each one is looked up to confirm it resolves to a registered record, at Crossref or, for preprints, at doi.org. A DOI that does not resolve is shown as unconfirmed rather than dropped, so you can see which entries to treat with care.
Metadata comes from the source
Titles, authors and years are reproduced as the databases hold them, and journal names are taken from the publisher's registered record at Crossref. Where a record is incomplete, you see the gap instead of a plausible guess.
You read before you pay
The free preview includes the field overview and three of the papers in full. If the quality is not there, you have lost nothing but a few minutes. Refunds are covered in the terms.
What it is not
It is a starting point, not a systematic review. It will miss papers a specialist in your sub-field would know. Treat the list as a first week of reading, then follow the citations yourself.

Why this exists

I spent the first months of my PhD reading more or less at random. I would find a paper, follow a citation, find another, and three weeks later I still could not have told you which five papers mattered most or what order they went in.

What I wanted was the thing a good supervisor gives you in the first week: a short list, in order, with a sentence about why each one is on it. I built that, and it now takes a few minutes instead of a few months.

Written by the person who built CiteStart, still a PhD student.

Price

One report, one price.

€0Free preview
The field overview and three of the papers in full, with their summaries and why each one is worth reading. Enough to judge the rest.
€15Full report
Every paper in your list, the complete reading order, trends, gaps, open questions, methods and datasets, keywords and next steps, plus the PDF. Paid once for that report, no subscription and no account.

Payment is handled by LemonSqueezy, who are the merchant of record. We never see your card details. If a report comes out badly, write to support@citestart.com and the refund policy applies.

Questions

How is this different from asking ChatGPT for papers?
A language model asked for citations will often produce ones that do not exist, with plausible authors and a DOI that resolves to nothing. Here the papers are retrieved from academic databases first, and each DOI is checked against Crossref. The model writes the summaries and the ordering; it never supplies the papers.
Which fields does it work for?
Anything indexed by Semantic Scholar and OpenAlex, which is most of the published literature. Coverage is strongest in computer science, engineering, biomedical sciences, physics and the quantitative social sciences. Humanities coverage is thinner, because less of that literature is indexed with usable metadata.
How many papers do I get?
Up to fifteen, and usually ten or more. A narrow topic yields fewer, and when that happens the report says so on the page rather than padding the list with loosely related work. Fifteen is the ceiling because it is about a week of serious reading, and small enough that the ordering still means something.
Can I share it with my supervisor?
Yes. The PDF is meant for that. Several sections, the gaps and the open questions in particular, are more useful in a supervision meeting than on your own.
What if the report is bad?
Read the free preview before paying; that is what it is for. If you pay and the result is poor, email support and the refund policy in the terms applies.
How long is my report kept?
Thirty days, at a private link. There are no accounts, so that link is the only way back to it. Save it or download the PDF.