RSC values range from 0 to larger positive values. complexity of the regulatory dictionary. As a direct application we used this unified binding repertoire to Cbll1 annotate variant enhancer loci (VELs) from H3K4me1 mark in two CHIR-99021 trihydrochloride cancer cell lines (MCF-7, CRC) and observed enrichments of specific TFs involved in biological key functions to cancer development and proliferation. Those enrichments of TFs within VELs provide a direct annotation of non-coding regions detected in CHIR-99021 trihydrochloride cancer genomes. Finally, full access to this catalogue is available online together with the TFs enrichment analysis tool (http://tagc.univ-mrs.fr/remap/). == INTRODUCTION == Differences in gene expression programs are believed to play a major role in cell identity and phenotypic diversity in the human body. With the advances of next generation sequencing techniques it became possible to study the genome-wide occupancy maps of transcription factors (TFs) by chromatin immunoprecipitation followed by sequencing (ChIP-seq). The rapid accumulation of ChIP-seq results in data warehouses provides a unique resource of hundreds of occupancy maps. With the success of the Encyclopedia of DNA elements (ENCODE) project to identify all functional elements in the human genome, the description and annotation of TF binding sites (TFBS) entered a genome-wide era by integrating a hundred TFs. The extent to which the regulatory space is organized along the genome is only starting to unfold with large consortia studies (15) but remains largely matter of discoveries (e. g. super enhancers) (6, 7). Indeed, recent studies have uncovered hundreds of genomic loci that are co-occupied by multiple TFs in various cell types suggesting the importance and abundance of combinatorial regulation CHIR-99021 trihydrochloride in cells (2, 7, 8). So far, those diverse regulatory features generated from various studies have not been yet integrated to form a global map of regulatory elements. Here, we report the complex landscape of TFBS in the human genome. We have constructed a global map of regulatory elements by compiling the genomic localization of 132 different TFs across 83 different cell lines and tissue types based on 395 selected human public (non-ENCODE) data sets. The integration of the genome-wide TFBS allows for the construction of a catalogue ofcis-regulatory modules (CRMs) of variable complexity. Specifically, we report a complex map of TFBS increasing by 14% (+439 Mb, +993 421 regulatory features) the human genome regulatory search space compared to the ENCODE catalogue alone. Different studies have proposed to integrate various NGS ChIP-seq data sets but either from a workflow/platform approach (9) or from a quality assessment CHIR-99021 trihydrochloride perspective (10). We performed a detailed comparison of the TFs occupancy map generated from public data against the ENCODE TF catalogue. Both maps are examined at the levels of TFs, TFBS and CRMs scales; finally, both maps are merged allowing the creation of a large catalogue of complex organization of bound regions. In total, after including ENCODE TF data to complement public data, we examined 237 TFs across multiple cell types. This map has been compiled into public tracks in genome browsers allowing users to assess regulatory elements combined with genome annotations in their regions of interests. In addition , the catalogue has been compiled into flat files allowing further computational analyses and is available athttp://tagc.univ-mrs.fr/remap/. Finally, to demonstrate the usefulness of our approach we used this unified catalogue to annotate variant enhancer loci (VELs; H3K4me1 mark) from two cancer cell lines. Our TFs enrichment analyses within those variable regions reveal enrichments of specific TFs those functions are involved in cancer development and proliferation. The work presented here constitutes a solid unification of regulatory regions in the human genome using a systematic integration of public non-ENCODE and ENCODE data. Taken together with our TF enrichment tool it allows for a better annotation of enhancers. == MATERIALS AND METHODS == == Public/non-ENCODE data sets sources == Public ChIP-seq data sets were extracted from Gene Expression Omnibus (GEO) and ArrayExpress (AE) databases. For GEO, the query(chip seq OR chipseq OR chip sequencing) AND Genome binding/occupancy profiling by high throughput sequencing AND homo sapiens[organism] AND NOT ENCODE[project]was used to return a.