SA-‐Mot: a web server for the iden fica on of mo fs of interest

Transcription

SA-‐Mot: a web server for the iden fica on of mo fs of interest
SA-­‐Mot: a web server for the iden5fica5on of mo5fs of interest extracted from protein loops Leslie REGAD MTi, UMR-­‐S973 Inserm, Univ Paris Diderot JOBIM, Rennes, 03 juillet 2012 Protein loops: •  link between regular secondary structures: α-­‐helix and β-­‐strand •  involved in the ligand specificity and protein func=on •  variable regions in terms of structure and sequence •  previous studies of SHORT loops: o  extrac=on of small structural mo=fs from short loops (Hutchinson et al., 1994) o  short loop classifica=on (Wojick et al., 1999, Fernandez et al., 2004) Extrac=on of some structural mo=fs from SHORT loops SA-­‐Mot Web server protocole hUps://sa-­‐mot.m=.univ-­‐paris-­‐diderot.fr •  Analysis of all loops of given protein Protein target: Topo II complexed with ANP Loop informa:ons SA-­‐Mot SA-­‐Mot mo=f database SA-­‐Mot motif database •  Contains important loop mo=fs : –  For the protein structure –  For the protein func=on •  Extrac=on mo=fs from loops: –  Based on o  Structural alphabet HMM-­‐SA (Camproux et al., 2004) o  Structural word no=on (Regad et al., 2010) –  Without superimposi=on of loop structures Presentation of HMM-­‐SA Camproux et al., J Mol Biol, 2004 •  Use HMM-­‐SA (Hidden Markov Model – Structural Alphabet) –  Classifica=on of all 4-­‐residue fragments extracted from a large set of protein structures •  Based on the geometry of 4-­‐residue fragments •  Used hidden Markov model & A L
G
G
G
G
& AL
& A L
& A L
Simplification of a protein structure using HMM-­‐SA Camproux et al., J mol Biol, 2004 Development of the STRUCTURAL WORD no=on Extraction of structural motif HMM-­‐SA Regad et al., CSDA, 2008 •  Extrac=on of 7-­‐residue structural mo=fs using HMM-­‐SA –  4 SL-­‐words = 7-­‐residue fragments 1gpw_B Pos: 7-­‐13
VDVKNGK 1sfi_A Pos: 28-­‐34 AEYHNTQ Extraction of structural motifs using HMM-­‐SA Regad et al., BMC Bioinfo, 2010 YUOD: 183 7-­‐AA fragments • 
DRPI: 389 7-­‐AA fragments PZCD : 983 7-­‐AA fragments Recurrent structural words: cluster of 7-­‐AA fragments with similar geometry (RMSd = 0.85 Å) • 
Extrac=on of structural word !
rapid extrac=on of structural mo=fs from protein loops without superimposi=on of protein structures SA-­‐Mot database Regad et al., BMC Bioinfo, 2010 • 
4,911 protein structures –  Non redundant (< 50% of iden=ty sequence) –  High resolu=on (< 2.5 Å) –  Classified in SCOP classifica=on Loop extrac=on   90, 811 loops of 7 to 35 residues   25,305 ≠ structural words seen between 1 to 1234 =mes = 238,158 7-­‐residue fragments Word extrac=on 2-­‐ Different type of SA-­‐Mot motifs Regad et al., BMC Bioinfo, 2010 • 
Occurrence: –  Rare words: link to uncertain region –  Recurrent words: regular structures •  Weak RMSD •  Strong AA-­‐score •  Over-­‐representated in loop dataset: non-­‐random mo@fs 
Regular structures with conserved structures and amino acid specifici@es PZCD: 983 7-­‐AA fragments DRPI: 389 7-­‐AA fragments Link between structural words and binding sites Regad et al., BMC Bioinfo, 2011 •  Loops ooen involved in protein func=on •  Over-­‐representa=on of word in each SCOP superfamilies Extrac=on of specific words for protein func=on 3-­‐ Quantification of the specificity of a motif for a function Superfamily : 52540 1gky 1ex7 1ex6 Superfamily: 81301 1no5 1wot 2i9g Simplifica=on of protein structures using HMM-­‐SA …KZaaaAZILPMMNDYUODNH…!
…CGBQAWZILPNNTRYUODNL…!
…SVZWaaZIMPNXTFYUODXX…!
…ZCGBQLNTMXJUBQKUOZILP…!
…JFCGBQSJMHVQLLQKUXMMN…!
…VOCGBQGIHBAABBQKUQXYZ…!
Computa=on of over-­‐represnta=on of 4-­‐SL words in SCOP superfamily Words Sf 52540 Sf 81301 ZILP OR OR SPECIFIC OF 2 SUPERFAMILIES YUOD OR NS SPECIFIC OF 1 SUPERFAMILY CGBQ NS OR SPECIFIC OF 1 SUPERFAMILY Other SA-­‐Mot motif types Regad et al., BMC Bioinfo, 2011 • 
Words specific to a lot of superfamilies • 
Word specific to one superfamily UBIQUITOUS WORDS important for protein structure CANDIDATE FUNCTIONAL WORDS important for protein func=on DODQ EIJU YUOD Ca2+-binding site
RUDO ATP/GTP-­‐binding site NAD(P)-­‐binding site SAM/SAH-­‐binding site SA-­‐Mot Input Regad et al., NAR, 2011 •  Our own pdb files •  Pdb code (4 characters) SA-­‐Mot outputs of 2rhmB Encoding into HMM-­‐SA SA-­‐Mot outputs of 2rhmB Encoding into HMM-­‐SA SA-­‐Mot mo=f database Word Occ Extrac=on of SA-­‐Mot mo=fs OR_l score RMSd AA Score OR_sf score type HBBQ 854 34 0.61 12 34 Ubi PKCD 0.4 0.54 2 0.7 2.5 0.55 23 120 56 YUOD 183 Fct SA-­‐Mot outputs of 2rhmB SA-­‐Mot outputs of 2rhmB SA-­‐Mot outputs of 2rhmB SA-­‐Mot outputs of 2rhmB SA-­‐Mot outputs of 2rhmB Exemple: search of important regions in uncharacterized protein •  puta=ve kinase from Chloroflexus auran:acus J-­‐10-­‐fl (pdb code 2rhm) SA-­‐Mot Exemple: search of important regions in uncharacterized protein Flexible regions Exemple: search of important regions in uncharacterized protein Over-­‐represented in ATP-­‐binding protein Over-­‐represented YVTN Repeat like Over-­‐represented Trypsin like Serine Protease Candidate func=onal sites ATP-­‐binding site Flexible regions Exemple: search of important regions in uncharacterized protein •  Iden=fica=on of structural mo=fs of interest –  For protein structure –  For protein func=on •  Puta=ve func=onal sites •  ATP-­‐binding site  in agreement with kinase ac=vity  no found with other predic=on mo=f server Exemple: Analysis of binding sites of topoisomerases of type II •  We have 29 structures : –  ≠ families : gyrase, Topo VI, Topo IV, Topo II Topo II complexed with ANP Topo VI complexed with ADP DNA gyrase complexed with ANP  Search of mo=fs of interest in these protein structures  Comparison of binding sites Exemple: Analysis of binding sites of topoisomerases of type II • 
DNA gyrase of E. coli (pdb code 1eij) structural mo@fs  Conserved region in terms of structure and sequence ATPase domain of DNA topoisomerase II NADP-­‐binding Rossman fold domain Immunoglobulin Rare region  Linked to flexible region Exemple: Analysis of binding sites of topoisomerases of type II • 
Structural word PSYR: seen in 20 structures –  Topo II do not contain this mo=f • 
Structural word ZQXU: seen in 4 Gyrases of E. coli • 
Structural word RNHB: seen only in 1eij protein Exemple: Analysis of binding sites of topoisomerases of type II • 
Structural word PSYR: seen in 20 structures –  Topo II do not contain this mo=f • 
Structural word ZQXU: seen in 4 Gyrases of E. coli • 
Structural word RNHB: seen only in 1eij protein DNA gyrase complexed with ANP Informa=on about the SPECIFICITY of binding Topo II complexed with ADP Conclusion SA – Mot: web server o  hUps://sa-­‐mot.m=.univ-­‐paris-­‐diderot.fr o  Extrac=on of strutural mo=fs of interest from protein loops without superimposi=on of protein structures o  could be used to   analysis of loop structures   comparison of loop structures  binding site  Analysis of protein deforma=on •  Perspec=ves: o  Search of degenerated words o  Generalize the approach to regular secondary structures Acknowledgement •  AC Camproux, MTi, Inserm UMR-­‐S973, Univ Paris Diderot •  C. Geneix, MTi, Inserm UMR-­‐S973, Univ Paris Diderot •  A. Saladin, MTi, Inserm UMR-­‐S973, Univ Paris Diderot •  J. Maupe=t, MTi, Inserm UMR-­‐S973, Univ Paris Diderot •  J. Mar=n, IBCP, CNRS, UMR 5086, Univ Lyon I •  G. Nuel, MAP5 -­‐ UMR CNRS 8145, Univ Paris Descartes •  C. Mayer, Ins=tut Pasteur, Univ Paris Diderot Thank you for your aUen=on !!!!!!!!! 

Documents pareils