9.3. PDB Fetchers — MDAnalysis.fetch.pdb

This suite of functions download structure files from the Research Collaboratory for Structural Bioinformatics (RCSB) Protein Data Batabank (PDB).

9.3.1. Variables

MDAnalysis.fetch.pdb.DEFAULT_CACHE_NAME_DOWNLOADER = 'MDAnalysis_pdbs'

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to sys.getdefaultencoding(). errors defaults to ‘strict’.

9.3.2. Functions

MDAnalysis.fetch.pdb.from_PDB(pdb_ids, cache_path=None, progressbar=False, file_format='cif.gz')[source]

Download one or more PDB files from the RCSB Protein Data Bank and cache them locally.

Given one or multiple PDB IDs, downloads the corresponding structure files format and stores them in a local cache directory. If files are cached on disk, from_PDB will skip the download and use the cached version instead.

Returns the path(s) as a Path to the downloaded file(s).

Parameters:
  • pdb_ids (str or sequence of str) – A single PDB ID as a string, or a sequence of PDB IDs to fetch.

  • cache_path (str or pathlib.Path) – Directory where downloaded file(s) will be cached. The default None argument uses the pooch default cache with project name DEFAULT_CACHE_NAME_DOWNLOADER.

  • file_format (str) – The file extension/format to download (e.g., “cif”, “pdb”). See the Notes section below for a list of all supported file formats.

  • progressbar (bool) – If True, display a progress bar during file downloads. Default is False.

Returns:

The path(s) to the downloaded file(s). Returns a single Path if a single pdb id is given, or a list of Path if multiple pdb ids are provided.

Return type:

Path or list of Path

Raises:

Notes

This function uses the RCSB File Download Services for directly downloading structure files via https.

The RCSB currently provides data in 'cif' , 'cif.gz' , 'bcif' , 'bcif.gz' , 'xml' , 'xml.gz' , 'pdb' , 'pdb.gz', 'pdb1', 'pdb1.gz' file formats and can therefore be downloaded. Not all of these formats can be currently read with MDAnalysis.

Caching, controlled by the cache_path parameter, is handled internally by pooch. The default cache name is taken from DEFAULT_CACHE_NAME_DOWNLOADER. To clear cache (and subsequently force re-fetching), it is required to delete the cache folder as specified by cache_path.

Examples

Download a single PDB file:

>>> mda.fetch.from_PDB("1AKE", file_format="cif")
'./MDAnalysis_pdbs/1AKE.cif'

Download multiple PDB files with a progress bar:

>>> mda.fetch.from_PDB(["1AKE", "4BWZ"], progressbar=True)
['./MDAnalysis_pdbs/1AKE.pdb.gz', './MDAnalysis_pdbs/4BWZ.pdb.gz']

Download a single PDB file and convert it to a universe:

>>> mda.Universe(mda.fetch.from_PDB("1AKE"), file_format="pdb.gz")
<Universe with 3816 atoms>

Download multiple PDB files and convert each of them into a universe:

>>> [mda.Universe(pdb) for pdb in mda.fetch.from_PDB(["1AKE", "4BWZ"], progressbar=True)]
[<Universe with 3816 atoms>, <Universe with 2824 atoms>]

New in version 2.11.0.

MDAnalysis.fetch.pdb.from_ALPHAFOLD(id, cache_path=None, progressbar=False, file_format='cif')[source]

Download one or more AlphaFold structure files and cache them locally.

Given one AlphaFold ID, downloads the corresponding structure file in the specified format and stores it in a local cache directory. If files are cached on disk, from_ALPHAFOLD will skip the download and use the cached version instead.

Returns the path(s) as a Path to the downloaded file(s).

Parameters:
  • id (str) – A single AlphaFold ID as a string.

  • cache_path (str or pathlib.Path) – Directory where downloaded file(s) will be cached. The default None argument uses the pooch default cache with project name DEFAULT_CACHE_NAME_DOWNLOADER.

  • file_format (str) – The file extension/format to download (e.g., “cif”, “pdb”). See the Notes section below for a list of all supported file formats.

  • progressbar (bool) – If True, display a progress bar during file downloads. Default is False.

Returns:

The path(s) to the downloaded file(s). Returns a single Path for the downloaded AlphaFold file.

Return type:

Path or list of Path

Raises:

Notes

This function uses the AlphaFold API for directly downloading structure files via HTTP GET.

AlphaFold currently provides data in 'cif', 'pdb', and 'bcif' file formats and can therefore be downloaded. Not all of these formats can be currently read with MDAnalysis.

At the current moment, this function only supports downloading a single AlphaFold ID at a time unlike in from_PDB() which can download multiple PDB IDs at once.

Additionally, there is currently no support for downloading prior versions of AlphaFold predictions. The AlphaFold API only provides the latest version of the prediction for a given ID. For more detailed control, it is recommended to browse AlphaFold manually.

Caching, controlled by the cache_path parameter, is handled internally by pooch. The default cache name is taken from DEFAULT_CACHE_NAME_DOWNLOADER. To clear cache (and subsequently force re-fetching), it is required to delete the cache folder as specified by cache_path.

Examples

Download a single AlphaFold file:

>>> from_ALPHAFOLD("Q9I1F6", file_format="cif")
'./MDAnalysis_pdbs/AF-Q9I1F6-F1-model_v6.cif'

Download a single AlphaFold file and convert it to a universe:

>>> mda.Universe(from_ALPHAFOLD("Q9I1F6"), files_format="pdb")
<Universe with 2608 atoms>

New in version 2.11.0.