Sequence search

Guide to searching the MGnify protein databases with HMMER
Author
Affiliation
Published

September 23, 2026

MGnify protein sequence searches are provided by the EMBL-EBI HMMER web HMMER web service. Use phmmer to search a protein query sequence in the MGnify Proteins Database.

Open HMMER

Choosing a MGnify30 database

MGnify30 contains representative sequences from MGnify protein clusters formed at 30% sequence identity. HMMER provides three subsets for different search goals:

Figure 3: MGnify30-C2 selected as the HMMER sequence database, with the default search cut-offs.

Non-singletons (MGnify30-C2)

The non-singletons subset was created by extracting clusters that have at least two members, thereby excluding a majority of the representatives. This subset contains 128,674,267 cluster representatives.

Larger non-singletons (MGnify30-C5-FL)

The larger non-singletons subset was created by extracting clusters that have at least five members, including at least one member that is predicted to be a full-length sequence. This subset therefore reduces the space even further, and guarantees at least one full-length sequence for increased confidence. This subset contains 20,911,652 cluster representatives.

Pfam-poor non-singletons (MGnify30-C5-PPfam)

The Pfam-poor non-singletons subset was created by extracting clusters that have at least five members, and where at least 90% of the cluster members do not have a Pfam accession. This subset similarly reduces the space significantly, and contains clusters for which function is relatively unknown. This subset contains 18,155,509 cluster representatives.

See the HMMER documentation for the current list and definitions of target databases.

Recording and interpreting results

The target databases are updated periodically. Open Search Details above the HMMER results table to find the target database name, release date and other search settings. Record these details with the query sequence when a search needs to be reproduced or reported.

Figure 4: The Search Details panel records the phmmer command, MGnify30 database release and query sequence.

HMMER’s results documentation explains the results table, domain graphics, alignments, scores and E-values.

Downloading MGnify protein data

MGnify protein releases and supporting files are also available from the MGnify FTP server. For questions or feedback, please contact the MGnify team.

Citation

BibTeX citation:
@online{2026,
  author = {, MGnify},
  title = {Sequence Search},
  date = {2026-09-23},
  url = {https://docs.mgnify.org/src/docs/mgnify-proteins-sequence-search.html},
  langid = {en}
}
For attribution, please cite this work as:
MGnify. 2026. “Sequence Search.” September 23. https://docs.mgnify.org/src/docs/mgnify-proteins-sequence-search.html.