idx_uniq |
unique index for the motif within the compendium |
motif_name |
unique name for the motif within the compendium, where the prefix matches the idx_uniq column and suffix matches the annotation column |
motif_name_safe |
computationally safe version of motif_name without slashes or pipe |
class |
class representing motif contribution to accessibility (Activating or Repressive) |
total_instances_across_celltypes |
total number of motif instances across all 189 cell types used for downstream modeling analysis |
cwm_logo_fwd |
name of the contribution weight matrix (CWM) logo (forward) PNG file (see Table S14) |
cwm_logo_rev |
name of the contribution weight matrix (CWM) logo (reverse) PNG file (see Table S14) |
query_consensus |
consensus sequence for trimmed CWM |
annotation |
granular motif annotation, in the form "<family>:<TF>#<index>” when a specific TF is unambiguous, or "<family>:<subfamily>/<alternative_subfamily>#<index>” where families and subfamilies of TFs could recognize the same motif, following the naming convention from the Vierstra v2.1 motif clustering. Index is used to distinguish variants of the same motif, typically representing subtle variations in nucleotide preferences or flanks. For composite motifs, constituents are separated by underscores. |
annotation_broad |
broad motif annotation, used throughout analyses and figures to group motifs together |
category |
motif category (base, basewithflanks, homocomposite, heterocomposite, partial, repeat, unresolved) |
best_match_TF |
TF associated with best matching known motif in databases (see Methods) |
best_match_motif |
best matching known motif in databases |
best_match_TOMTOM_qval |
q-value from TOMTOM representing similarity between best matching known motif |
best_match_TFs_in_family |
other TFs in the same family as bestmatchTF, obtained from HOCOMOCO |
best_match_TFs_in_Vierstra_archetype |
other TFs with motifs in the same cluster as the best matching known motif, according to the Vierstra v2.1 motif clustering |
total_seqlets_across_celltypes |
total number of seqlets (short loci used by TF-MoDISco to construct motifs) across all individual motifs which were aggregated into this compendium motif |
merged_pattern |
name of the cluster of per-cell-type motifs which corresponds to this compendium motif; can mapped back to the constituent motifs |
n_constituent_motifs |
number of motifs aggregated into this compendium motif |
n_constituent_celltypes |
number of cell types where at least one motif learned in that cell type (through TF-MoDISco) was aggregated into this compendium motif |
constituent_organs |
number of organs where at least one motif learned in a cell type in that organ (through TF-MoDISco) was aggregated into this compendium motif |
constituent_celltypes |
L2_annot for the cell types where at least one motif learned in that cell type was aggregated into this compendium motif |