Key responsibilities:
-
Work with senior team members to identify opportunities for simplification, automation, and continuous improvement in operational data processes, including modernization of data curation workflows
-
Support the development and improvement of automated workflows for data processing, quality control, recurring dataset updates, and AI-enabled process improvement
-
Collaborate with cross-functional teams to curate, maintain, and improve key oncology reference datasets, with an emphasis on quality, correctness, reproducibility, and operational reliability
-
Contribute to reference dataset lifecycle management, including documentation, versioning, and traceable update processes
-
Assess dataset structure, integrity, and fitness for downstream scientific use, including attention to identifiers, primary keys, missing values, consistency, and related data quality issues, and help resolve data issues in collaboration with relevant partners
-
Contribute to best practices for reproducible workflows, data handling, and scalable dataset operations
-
Help keep oncology reference datasets current, usable, and responsive to evolving scientific and organizational priorities
Essential Requirements:
-
Degree in bioinformatics, computational biology, data science, computer science, or a related field, or equivalent relevant experience
-
Experience working with omics datasets in a research environment
-
Proficiency in one or more programming languages used in data science or bioinformatics, such as Python or R
-
Experience with data wrangling, data quality control, and reproducible analysis workflows
-
Familiarity with workflow automation and scripting for recurring data processing tasks
-
Familiarity with database principles and sound data modeling practices, including normalization, primary keys, joins, and thoughtful handling of missing values and data integrity issues
-
Knowledge of version control and Unix / Linux-based working environments
-
Strong organizational skills and attention to detail, with a demonstrated interest in operational excellence
-
Ability to learn quickly, manage multiple tasks, and work effectively in a collaborative environment
-
Excellent written and verbal communication skills, and the ability to work effectively across technical and scientific teams
Desirable Requirements:
-
Experience with oncology datasets, translational research data, or biomedical reference data resources
-
Familiarity with FAIR data principles, metadata curation, or scientific data stewardship
-
Experience maintaining reference datasets or supporting data products used by multiple stakeholders
-
Experience collaborating with experimental scientists or lab-based partners
Novartis Compensation and Benefit Summary:
The salary for this position is expected to range between $103,600 and $192,400 per year.
The final salary offered is determined based on factors like, but not limited to, relevant skills and
experience, and upon joining Novartis will be reviewed periodically. Novartis may change the published
salary range based on company and market factors.
Your compensation will include a performance-based cash incentive and, depending on the level of the
role, eligibility to be considered for annual equity awards.
US-based eligible employees will receive a comprehensive benefits package that includes health, life and
disability benefits, a 401(k) with company contribution and match, and a variety of other benefits. In addition, employees are eligible for a generous time off package including vacation, personal days, holidays and other leaves.