Joint modeling of DNA sequence and physical properties to improve eukaryotic promoter recognition


  • U. Ohler
  • H Niemann
  • G.C. Liao
  • G.M. Rubin


  • Bioinformatics


  • Bioinformatics 17 Suppl 1: S199-206


  • We present an approach to integrate physical properties of DNA, such as DNA bendability or GC content, into our probabilistic promoter recognition system McPROMOTER. In the new model, a promoter is represented as a sequence of consecutive segments represented by joint likelihoods for DNA sequence and profiles of physical properties. Sequence likelihoods are modeled with interpolated Markov chains, physical properties with Gaussian distributions. The background uses two joint sequence/profile models for coding and non-coding sequences, each consisting of a mixture of a sense and an anti-sense submodel. On a large Drosophila test set, we achieved a reduction of about 30% of false positives when compared with a model solely based on sequence likelihoods.