This repository contains the workflow used for selecting, processing, and post-processing public proteomics datasets from the PRIDE database for integration into the Open Targets Platform.
The complete list of datasets in the PRIDE database was obtained using the PRIDE API.
Run:
Download_PRIDE_Datasets.pyThe resulting TSV file was manually filtered according to the dataset selection criteria:
Annotate sample metadata according to the SDRF guidelines.
SDRF specification
Annotated SDRF files
Process raw files using MaxQuant.
Important: The assay name field in the SDRF file must exactly match the Experiment field in MaxQuant.
Process raw files using DIA-NN.
Run:
OpenTargets_dataset_Summary_reportfile.pyRequired input files
proteinGroups.txt- SDRF file
Run:
OpenTargets_DIA_Summary_report.pyRequired input files
report.tsv- SDRF file
| Data Type | Script |
|---|---|
| DDA / TMT / iTRAQ | OpenTargets_dataset_Summary_reportfile.py |
| DIA | OpenTargets_DIA_Summary_report.py |
Raw and post-processed results are available from the PRIDE FTP server:
https://ftp.pride.ebi.ac.uk/pub/databases/pride/resources/proteomes/otargets/