# cath_v4_5_0.prerelease.CTA_results.tsv.gz

For the CATH v4.5.0 prerelease.

## What this data represents

All High and Medium consensus domains identified and segmented by the CTA nextflow pipeline from 10 million single chain AFDB v6 models (batch 1).

Gzipped tab-separated text. Header on line 1. One domain per row.

## How the data was generated

Over 65 million new or remodelled protein structures were downloaded from the AFDB/EBI database. These were batched into an intial batch (1) of 10 million protein chains. 

For each chain, the Uniprot accession was extracted from the name and queried against the Uniprot API for taxonomic information. Chains were then filtered to remove very short sequences.

The surviving chains were then submitted to the three domain segmentation algorithms: Chainsaw, Merizo and Unidoc. Considering algorithmic agreement and overlap, the resulting domains were then 

organised into consensus domains labelled H for high (all three methods agree) or M for medium (2/3 methods agree). Low consensus domains were discarded.

The domain boundaries are displayed (chopping) as well as STRIDE secondary structure assignments. Additionally packing density and normed radius of gyration are calculated using CATH-AF-CLI programs.

Average plddt is calculated from the plddt scores in the Bfac column and two further quality measures; domqual (from the Merizo-search pipeline) and dom single domain identifier are derived.

In addition, foldseek is used to query each domain against a CATH v4.4.0 s95 reference database and report a topology (T) match or a homology (H) match and associated foldseek scores.

## Columns

| Column | Meaning |
|---|---|
| uniprot_id | domain identifier |
| md5_domain | unique md5 hash for this sequence |  
| consensus_level | High (H) or Medium (M) consensus | 
| chopping | the domain boundaries |
| nres_domain | number of amino acids in the domain |
| num_segments | number of different segments in the domain |
| num_helix_strand_turn | number of all secondary structure elements in the domain |
| num_helix | number of helices in the domain |
| num_strand | number of beta strands in the domain |
| num_helix_strand | combined number of helices and strands in the domain |
| num_turn | number of turns in the domain |
| packing_density | local compactness or contact measure of the domain |
| normed_radius_gyration | a measure of atom spread around the centre of mass |
| avg_plddt | the mean plddt across the domain residues |
| proteome_id | Uniprot Proteome identifier |
| tax_common_name | Uniprot Organism name field |
| tax_scientific_name | Uniprot Organism name field; includes the scientific name |
| tax_lineage | Uniprot Taxonomic lineage |
| domqual | a measure of domain quality 0-1 score |
| dom_single_domain | indicates whether a domain is truly single (T/F) |
| foldseek_match_id | the CATH domain identifier |
| foldseek_evalue | foldseek E-value |
| foldseek_tmscore | the greater of Foldseek query or target TM |
| cath_label | CATH domain classification |
| foldseek_match_type | topology or homology level match |
| foldseek_query_cov | Foldseek query coverage |
| foldseek_target_cov | Foldseek target coverage |
| Q_score | a calculation of domain quality that includes the foldseek e-value |

## Shape

17,539,030 data rows.
