← All services
Service · Bioinformatics

Bioinformatics Solutions

As India's official Element Biosciences distributor, we deliver end-to-end B2B genomic reporting and custom bioinformatics, with live deployments at multiple labs and wellness chains.

Products & platforms

Production-grade consumer and preventive genomic reporting, plus deployable GUI tools and pipeline platforms built on Django, Docker and AWS.

Pipelines at scale

From microarray genotyping (IDAT-to-Report) to end-to-end NGS workflows (FASTQ-to-VCF-to-Report), delivering scalable bioinformatics pipelines supporting 700+ wellness traits, 1,800+ carrier disorders, 40+ hereditary cancer conditions, and pharmacogenomic reporting for 100+ medications.

Data residency

The production array pipeline runs on an in-house server in Gurugram, where only finished reports cross the facility boundary. Customer-premise, self-managed AWS and managed AWS HealthOmics deployments are supported as well, with encryption, access controls and HIPAA-grade compliance throughout.

Deployable platforms

Django and Docker GUI tools (interactive dashboards, report viewers and pipeline UIs), deployable on-premise or on AWS EC2, with live platforms running at multiple labs and wellness chains.

ElemBio support & R&D

NGS troubleshooting and analysis support for labs, plus advanced R&D in CGP, ctDNA, scRNA-Seq and multi-omics through Geniquely.

From raw data to a variant call

The first half of the work is deliberately not proprietary. Genique runs published best practice on open-source tools, because a pipeline a laboratory cannot audit is a pipeline it cannot defend at an accreditation assessment.

  • Whole genome, three stages and nine steps: adapter trimming and QC with fastp, alignment with BWA-MEM2, sorting and duplicate marking with samtools and Picard, base quality score recalibration, per-sample calling with GATK HaplotypeCaller in GVCF mode, joint genotyping, variant recalibration, genotype refinement, and functional annotation with VEP.
  • Array, three steps: genotype calling from raw IDAT with GenCall, sample and variant QC with PLINK and bcftools, and a strand and build check, ending at a QC-passed VCF.
  • Depth: 30x mean coverage is the standard for clinical germline whole genome work, consistent with ACMG and CAP guidance, with 90 to 95 per cent of bases at 20x or above. Below 15x, the false-negative rate on heterozygous calls makes a clinical report difficult to defend.
  • Scale: a 30x whole genome sample is 90 to 120 GB of raw data against 20 to 60 MB for a GSA array sample, which is why the two are costed and scheduled as separate services rather than as one.
  • Orchestration: Nextflow, including the nf-core sarek implementation, or Cromwell, with Slurm scheduling across concurrent samples.
  • Volume: more than 18,000 samples processed to date across whole exome, whole genome and GSA array work, in business-to-business partner programmes.

What each data source can report

Array and whole genome are not two prices for the same answer. Four of the eight report types are limited or unavailable from array data, and the reason is structural rather than a matter of analysis quality.

  • Both, in full: wellness and ancestry.
  • Whole genome only: HLA typing, because that region is the most polymorphic in the genome and cannot be resolved from fixed chip positions, and secondary findings, which need full gene coverage.
  • Whole genome in full, array in part: hereditary disease and cancer risk, carrier status, pharmacogenomics and polygenic scores. An array genotypes known common positions, so rare and novel variants, indels and copy-number changes are not detected.
  • Two loci that need copy-number-aware calling whichever the source: SMN1, where the near-identical SMN2 paralog leads standard callers to under-call it, and CYP2D6, where the CYP2D7 pseudogene does the same. Both need dedicated callers rather than a general variant caller.
  • The array track adds steps before any report is written: imputation against the 1000 Genomes reference, PCA projection onto HGDP and 1000 Genomes for ancestry, and polygenic scoring.

Eight report types from one variant call file

The interpretation layer is where the Genique work sits, and it is built on published frameworks rather than on a private definition of risk.

  • Hereditary disease and cancer risk: the ACMG and AMP five-tier classification, applied to the ordered gene list. Only Pathogenic and Likely Pathogenic variants are reported for the indication. Variants of uncertain significance are disclosed and not actioned.
  • Carrier status: condition inclusion follows ACMG and ACOG criteria on carrier frequency, phenotype severity and detectability, with reproductive risk assessed jointly once both partners are screened.
  • Secondary findings: the ACMG SF v3.2 list, 81 genes covering 27 cancer, 45 cardiovascular and 9 metabolic and other conditions, reported independent of the indication, with opt-out as standard practice.
  • Pharmacogenomics and HLA: star-allele calling across CYP2D6, CYP2C19, CYP2C9, CYP3A5, VKORC1, TPMT, DPYD, SLCO1B1 and UGT1A1, against PharmVar allele definitions and CPIC dosing guidance, and allele-level HLA typing against the IMGT and HLA reference.
  • Polygenic scores: harmonised scorefiles from the PGS Catalog, scored with pgsc_calc and reported as a percentile within an ancestry-matched reference distribution. A percentile is a rank against a reference population and not a probability of illness. Each score carries an evidence tier: A where it is validated in four or more ancestry groups with a same-ancestry performance metric, B in three, C where it has not yet been evaluated in the reported ancestry.
  • Wellness and ancestry: a curated marker catalogue of 20 categories, 51 subcategories, 166 traits, 482 markers and 539 genes, which is one cut of the catalogue rather than the whole of it, and ancestry as statistical admixture from PCA projection onto HGDP and 1000 Genomes, with every reference population in the model disclosed and proportions summing to 100 per cent. Ancestry here is statistical similarity to reference genomes. It is not nationality, religion, caste or self-identified ethnicity.

Where the pipeline runs, and where the data stays

For most Indian customers this question is settled by data residency before it is settled by performance, so it is worth saying plainly which of the four deployments applies.

  • On premise at Genique: the production array pipeline runs containerised and version-pinned on an in-house server in Gurugram, with a local copy of every reference resource. Only finished PDF reports cross the facility boundary. Raw genotypes, imputed data and scores do not.
  • On premise at the customer: the binding constraints are alignment and variant calling across 32 or more threads, 32 GB or more of RAM for annotation and duplicate marking, and fast local NVMe scratch. Network-mounted scratch has collapsed throughput in our experience, so it is not the saving it looks like.
  • AWS, self-managed: S3 with Batch or ParallelCluster on EC2, and FSx for Lustre as a high-throughput scratch filesystem. Deployable in the Mumbai region, which keeps patient data in India.
  • AWS HealthOmics, managed: the least infrastructure overhead of the four. It is not currently available in the Mumbai region, and the nearest are Singapore, Tokyo and Seoul, which makes it a data-residency decision rather than a technical one.
  • Storage planning: roughly 100 to 120 GB of BAM per whole genome sample during processing and about 20 GB as CRAM on completion, plus a raw FASTQ archive kept so a sample can be re-analysed against updated references, which is an accreditation requirement rather than a capacity choice.

Interested in Bioinformatics Solutions?

Talk to our team about how it fits your needs.

Talk to us