SA-‐Mot: a web server for the iden fica on of mo fs of interest
Transcription
SA-‐Mot: a web server for the iden fica on of mo fs of interest
SA-‐Mot: a web server for the iden5fica5on of mo5fs of interest extracted from protein loops Leslie REGAD MTi, UMR-‐S973 Inserm, Univ Paris Diderot JOBIM, Rennes, 03 juillet 2012 Protein loops: • link between regular secondary structures: α-‐helix and β-‐strand • involved in the ligand specificity and protein func=on • variable regions in terms of structure and sequence • previous studies of SHORT loops: o extrac=on of small structural mo=fs from short loops (Hutchinson et al., 1994) o short loop classifica=on (Wojick et al., 1999, Fernandez et al., 2004) Extrac=on of some structural mo=fs from SHORT loops SA-‐Mot Web server protocole hUps://sa-‐mot.m=.univ-‐paris-‐diderot.fr • Analysis of all loops of given protein Protein target: Topo II complexed with ANP Loop informa:ons SA-‐Mot SA-‐Mot mo=f database SA-‐Mot motif database • Contains important loop mo=fs : – For the protein structure – For the protein func=on • Extrac=on mo=fs from loops: – Based on o Structural alphabet HMM-‐SA (Camproux et al., 2004) o Structural word no=on (Regad et al., 2010) – Without superimposi=on of loop structures Presentation of HMM-‐SA Camproux et al., J Mol Biol, 2004 • Use HMM-‐SA (Hidden Markov Model – Structural Alphabet) – Classifica=on of all 4-‐residue fragments extracted from a large set of protein structures • Based on the geometry of 4-‐residue fragments • Used hidden Markov model & A L G G G G & AL & A L & A L Simplification of a protein structure using HMM-‐SA Camproux et al., J mol Biol, 2004 Development of the STRUCTURAL WORD no=on Extraction of structural motif HMM-‐SA Regad et al., CSDA, 2008 • Extrac=on of 7-‐residue structural mo=fs using HMM-‐SA – 4 SL-‐words = 7-‐residue fragments 1gpw_B Pos: 7-‐13 VDVKNGK 1sfi_A Pos: 28-‐34 AEYHNTQ Extraction of structural motifs using HMM-‐SA Regad et al., BMC Bioinfo, 2010 YUOD: 183 7-‐AA fragments • DRPI: 389 7-‐AA fragments PZCD : 983 7-‐AA fragments Recurrent structural words: cluster of 7-‐AA fragments with similar geometry (RMSd = 0.85 Å) • Extrac=on of structural word ! rapid extrac=on of structural mo=fs from protein loops without superimposi=on of protein structures SA-‐Mot database Regad et al., BMC Bioinfo, 2010 • 4,911 protein structures – Non redundant (< 50% of iden=ty sequence) – High resolu=on (< 2.5 Å) – Classified in SCOP classifica=on Loop extrac=on 90, 811 loops of 7 to 35 residues 25,305 ≠ structural words seen between 1 to 1234 =mes = 238,158 7-‐residue fragments Word extrac=on 2-‐ Different type of SA-‐Mot motifs Regad et al., BMC Bioinfo, 2010 • Occurrence: – Rare words: link to uncertain region – Recurrent words: regular structures • Weak RMSD • Strong AA-‐score • Over-‐representated in loop dataset: non-‐random mo@fs Regular structures with conserved structures and amino acid specifici@es PZCD: 983 7-‐AA fragments DRPI: 389 7-‐AA fragments Link between structural words and binding sites Regad et al., BMC Bioinfo, 2011 • Loops ooen involved in protein func=on • Over-‐representa=on of word in each SCOP superfamilies Extrac=on of specific words for protein func=on 3-‐ Quantification of the specificity of a motif for a function Superfamily : 52540 1gky 1ex7 1ex6 Superfamily: 81301 1no5 1wot 2i9g Simplifica=on of protein structures using HMM-‐SA …KZaaaAZILPMMNDYUODNH…! …CGBQAWZILPNNTRYUODNL…! …SVZWaaZIMPNXTFYUODXX…! …ZCGBQLNTMXJUBQKUOZILP…! …JFCGBQSJMHVQLLQKUXMMN…! …VOCGBQGIHBAABBQKUQXYZ…! Computa=on of over-‐represnta=on of 4-‐SL words in SCOP superfamily Words Sf 52540 Sf 81301 ZILP OR OR SPECIFIC OF 2 SUPERFAMILIES YUOD OR NS SPECIFIC OF 1 SUPERFAMILY CGBQ NS OR SPECIFIC OF 1 SUPERFAMILY Other SA-‐Mot motif types Regad et al., BMC Bioinfo, 2011 • Words specific to a lot of superfamilies • Word specific to one superfamily UBIQUITOUS WORDS important for protein structure CANDIDATE FUNCTIONAL WORDS important for protein func=on DODQ EIJU YUOD Ca2+-binding site RUDO ATP/GTP-‐binding site NAD(P)-‐binding site SAM/SAH-‐binding site SA-‐Mot Input Regad et al., NAR, 2011 • Our own pdb files • Pdb code (4 characters) SA-‐Mot outputs of 2rhmB Encoding into HMM-‐SA SA-‐Mot outputs of 2rhmB Encoding into HMM-‐SA SA-‐Mot mo=f database Word Occ Extrac=on of SA-‐Mot mo=fs OR_l score RMSd AA Score OR_sf score type HBBQ 854 34 0.61 12 34 Ubi PKCD 0.4 0.54 2 0.7 2.5 0.55 23 120 56 YUOD 183 Fct SA-‐Mot outputs of 2rhmB SA-‐Mot outputs of 2rhmB SA-‐Mot outputs of 2rhmB SA-‐Mot outputs of 2rhmB SA-‐Mot outputs of 2rhmB Exemple: search of important regions in uncharacterized protein • puta=ve kinase from Chloroflexus auran:acus J-‐10-‐fl (pdb code 2rhm) SA-‐Mot Exemple: search of important regions in uncharacterized protein Flexible regions Exemple: search of important regions in uncharacterized protein Over-‐represented in ATP-‐binding protein Over-‐represented YVTN Repeat like Over-‐represented Trypsin like Serine Protease Candidate func=onal sites ATP-‐binding site Flexible regions Exemple: search of important regions in uncharacterized protein • Iden=fica=on of structural mo=fs of interest – For protein structure – For protein func=on • Puta=ve func=onal sites • ATP-‐binding site in agreement with kinase ac=vity no found with other predic=on mo=f server Exemple: Analysis of binding sites of topoisomerases of type II • We have 29 structures : – ≠ families : gyrase, Topo VI, Topo IV, Topo II Topo II complexed with ANP Topo VI complexed with ADP DNA gyrase complexed with ANP Search of mo=fs of interest in these protein structures Comparison of binding sites Exemple: Analysis of binding sites of topoisomerases of type II • DNA gyrase of E. coli (pdb code 1eij) structural mo@fs Conserved region in terms of structure and sequence ATPase domain of DNA topoisomerase II NADP-‐binding Rossman fold domain Immunoglobulin Rare region Linked to flexible region Exemple: Analysis of binding sites of topoisomerases of type II • Structural word PSYR: seen in 20 structures – Topo II do not contain this mo=f • Structural word ZQXU: seen in 4 Gyrases of E. coli • Structural word RNHB: seen only in 1eij protein Exemple: Analysis of binding sites of topoisomerases of type II • Structural word PSYR: seen in 20 structures – Topo II do not contain this mo=f • Structural word ZQXU: seen in 4 Gyrases of E. coli • Structural word RNHB: seen only in 1eij protein DNA gyrase complexed with ANP Informa=on about the SPECIFICITY of binding Topo II complexed with ADP Conclusion SA – Mot: web server o hUps://sa-‐mot.m=.univ-‐paris-‐diderot.fr o Extrac=on of strutural mo=fs of interest from protein loops without superimposi=on of protein structures o could be used to analysis of loop structures comparison of loop structures binding site Analysis of protein deforma=on • Perspec=ves: o Search of degenerated words o Generalize the approach to regular secondary structures Acknowledgement • AC Camproux, MTi, Inserm UMR-‐S973, Univ Paris Diderot • C. Geneix, MTi, Inserm UMR-‐S973, Univ Paris Diderot • A. Saladin, MTi, Inserm UMR-‐S973, Univ Paris Diderot • J. Maupe=t, MTi, Inserm UMR-‐S973, Univ Paris Diderot • J. Mar=n, IBCP, CNRS, UMR 5086, Univ Lyon I • G. Nuel, MAP5 -‐ UMR CNRS 8145, Univ Paris Descartes • C. Mayer, Ins=tut Pasteur, Univ Paris Diderot Thank you for your aUen=on !!!!!!!!!