Help
» Japanese
KEGG icon

KEGG Overview

1. Genomes to Biological Systems

The KEGG database has been developed as a computer model of the biological systems, such as the cell, the organism and the biosphere, enabling to uncover higher-level systemic functions and utilities from genomes, metagenomes and other molecular-level datasets. The KEGG model consists of KEGG pathway maps and other molecular networks of interactions, reactions and relations, which are manually created from published literature and linked to/from genomes through the KEGG Orthology (KO) system. In comparison to the AI model as shown below, the KEGG model is focused on organizing our knowledge in life sciences, and its methodology to uncover biological systems from genomes may have applications in human health and environmental sustainability. KEGG model

2. The KEGG Database

The KEGG model is realized by a collection of data objects for representation of both generic and human-specific biological information systems. They are implemented as 23 types of data objects stored in 13 databases. They are also broadly categorized into systems, genomic, chemical and health information, and are distinguished by color coding of web pages.

Category Database Data object Color
Systems
information
PATHWAY pathway KEGG pathway maps kegg3
BRITE brite BRITE hierarchies and tables
MODULE module KEGG modules
Genomic
information
KO ko KO functional orthologs kegg4
GENES <org>
ag
vg
vp
Genes in KEGG organisms
Addendum proteins
Viral genes
Viral mature peptides
kegg1
GENOME genome
vtax
vgenome
KEGG organisms
KEGG viruses
Pathogenic viruses
Chemical
information
COMPOUND compound Metabolites and other chemical substances kegg2
GLYCAN glycan Glycans
REACTION reaction
rclass
rmodule
Biochemical reactions
Reaction class
Reaction module
ENZYME enzyme Enzyme nomenclature
Health
information
NETWORK network
ntmap
variant
Disease-related network elements
Network variantion maps
Human gene variants
kegg5
DISEASE disease Human diseases
DRUG drug
dgroup
Drugs
Drug groups
Health information category integrated with drug labels is called KEGG MEDICUS

Data objects in KEGG are highly integrated with various types of links.
Data object links
Most of them are developed internally using published literature as the primary data source. Exceptions are genes taken from NCBI RefSeq and GenBank, genome and vtax taken from NCBI Genome and Taxonomy, and enzyme taken from ExplorEnz.

3. History of KEGG

The KEGG project was initiated in 1995. For the KEGG-original data objects, the identifier (database entry name) takes the form of a predefined prefix and a five-digit number (see: KEGG objects). The following table shows how KEGG databases have been expanded.

Release Database Data object identifier Remark
1995KEGG PATHWAYmap number
KEGG GENESlocus_tag / GeneID
KEGG ENZYMEEC number
KEGG COMPOUNDC number
1998KEGG REACTIONR number
2000KEGG GENOMEorganism code / T number
2002KEGG ORTHOLOGY  K numberOrtholog IDs in 2000
2003KEGG GLYCANG number
2004KEGG RPAIRRP numberDiscontinued in 2016
2005KEGG BRITEbr number
KEGG DRUGD number
2006KEGG MODULEM number
2008KEGG DISEASEH number
2010KEGG RCLASSRC number
KEGG EDRUGE numberRenamed to ENVIRON
2011KEGG ENVIRONE numberDiscontinued in 2021
2012KEGG RMODULERM number
2014KEGG DGROUPDG number
2017KEGG NETWORKN number / nt number
KEGG VARIANTGeneID+variant number

4. KEGG Molecular Networks

The most unique data object in KEGG is the molecular networks -- molecular interaction, reaction and relation networks representing systemic functions of the cell and the organism. Experimental knowledge on such systemic functions is captured from literature and organized in the following forms: These databases constitute the reference knowledge base for biological interpretation of genomes and high-throughput molecular datasets through the process of KEGG mapping (see: KEGG mapping).

In 1995 the concept of mapping was first introduced in KEGG for linking genomes to metabolic pathways (metabolic reconstruction) using the EC number. Once the EC numbers were assigned to enzyme genes in the genome, organism-specific pathways could be generated automatically by matching against the enzyme (EC number) networks of the KEGG reference metabolic pathways. The EC number is no longer used as an identifier in KEGG. The KEGG Orthology (KO) system is the basis for genome annotation and KEGG mapping.

Period Identifier Reference knowledge Assignment
1995-1999 EC number Metabolic pathways Domain based
2000-2002 Ortholog ID Metabolic and regulatory pathways Domain based
2003- KO Pathways and BRITE hierarchies Gene based

From a different perspective, individual instances of genes are grouped into KO entries representing functional orthologs in the molecular networks. There are two more types of such generalization in KEGG as shown below.

Network type Class Instance
All types KO (gene ortholog) Genes in KEGG GENES
Biochemical reaction RC (reaction class) Reactions in KEGG REACTION
Drug interaction DG (drug group) Drugs in KEGG DRUG

5. Network Variants

The KEGG database has been developed by focusing on conservation and variation of genes and genomes among different organisms. The reference datasets of KEGG pathway maps, BRITE hierarchies and KEGG modules have been developed with the concept of functional orthologs (KOs), so that KEGG pathway mapping and other procedures can be applied to any cellular organism.

KEGG variation

However, this generic approach is inadequate for understanding more detailed features caused by variations of genes and genomes within a species, especially for understanding disease related variations of human genes and genomes. KEGG NETWORK represents a renewed attempt by KEGG to capture knowledge on diseases and drugs in terms of network variants caused by not only gene variants, but also viruses and other factors.

References

  1. Kanehisa, M.; Toward pathway engineering: a new database of genetic and molecular pathways. Science & Technology Japan, No. 59, pp. 34-38 (1996). [pdf]
  2. Kanehisa, M.; A database for post-genome analysis. Trends Genet. 13, 375-376 (1997). [pubmed] [doi]
  3. Ogata, H., Goto, S., Sato, K., Fujibuchi, W., Bono, H., and Kanehisa, M.; KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 27, 29-34 (1999). [pubmed] [doi]
  4. Kanehisa, M. and Goto, S.; KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 28, 27-30 (2000). [pubmed] [doi]
  5. Kanehisa, M., Goto, S., Kawashima, S., and Nakaya, A.; The KEGG databases at GenomeNet. Nucleic Acids Res. 30, 42-46 (2002). [pubmed] [doi]
  6. Kanehisa, M., Goto, S., Kawashima, S., Okuno, Y., and Hattori, M.; The KEGG resources for deciphering the genome. Nucleic Acids Res. 32, D277-D280 (2004). [pubmed] [doi]
  7. Kanehisa, M., Goto, S., Hattori, M., Aoki-Kinoshita, K.F., Itoh, M., Kawashima, S., Katayama, T., Araki, M., and Hirakawa, M.; From genomics to chemical genomics: new developments in KEGG. Nucleic Acids Res. 34, D354-357 (2006). [pubmed] [doi]
  8. Kanehisa, M., Araki, M., Goto, S., Hattori, M., Hirakawa, M., Itoh, M., Katayama, T., Kawashima, S., Okuda, S., Tokimatsu, T., and Yamanishi, Y.; KEGG for linking genomes to life and the environment. Nucleic Acids Res. 36, D480-D484 (2008). [pubmed] [doi]
  9. Kanehisa, M., Goto, S., Furumichi, M., Tanabe, M., and Hirakawa, M.; KEGG for representation and analysis of molecular networks involving diseases and drugs. Nucleic Acids Res. 38, D355-D360 (2010). [pubmed] [doi]
  10. Kanehisa, M., Goto, S., Sato, Y., Furumichi, M., and Tanabe, M.; KEGG for integration and interpretation of large-scale molecular datasets. Nucleic Acids Res. 40, D109-D114 (2012). [pubmed] [doi]
  11. Kanehisa, M., Goto, S., Sato, Y., Kawashima, M., Furumichi, M., and Tanabe, M.; Data, information, knowledge and principle: back to metabolism in KEGG. Nucleic Acids Res. 42, D199–D205 (2014). [pubmed] [doi]
  12. Kanehisa, M., Sato, Y., Kawashima, M., Furumichi, M., and Tanabe, M.; KEGG as a reference resource for gene and protein annotation. Nucleic Acids Res. 44, D457-D462 (2016). [pubmed] [doi]
  13. Kanehisa, Furumichi, M., Tanabe, M., Sato, Y., and Morishima, K.; KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res. 45, D353-D361 (2017). [pubmed] [doi]
  14. Kanehisa, M., Sato, Y., Furumichi, M., Morishima, K., and Tanabe, M.; New approach for understanding genome variations in KEGG. Nucleic Acids Res. 47, D590-D595 (2019). [pubmed] [doi]
  15. Kanehisa, M; Toward understanding the origin and evolution of cellular organisms. Protein Sci. 28, 1947-1951 (2019). [pubmed] [doi]
  16. Kanehisa, M., Furumichi, M., Sato, Y., Ishiguro-Watanabe, M., and Tanabe, M.; KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 49, D545-D551 (2021). [pubmed] [doi]
  17. Kanehisa, M., Furumichi, M., Sato, Y., Kawashima, M. and Ishiguro-Watanabe, M.; KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 51, D587-D592 (2023). [pubmed] [doi]
  18. Kanehisa, M., Furumichi, M., Sato, Y., Matsuura, Y. and Ishiguro-Watanabe, M.; KEGG: biological systems database as a model of the real world. Nucleic Acids Res. 53, D672-D677 (2025). [pubmed] [doi]

Last updated: July 1, 2026

Copyright 1995-2026 Kanehisa Laboratories