See all roles

Data Scientist, Cancer Informatics and AI/ML, Remote, Grant Funded

Work from home Full-time role Hiring

Job Description

This position will support computational oncology and cancer informatics research initiatives focused on transforming complex clinical data into structured, actionable datasets for research, quality improvement, clinical trial identification, and care delivery optimization. The role will emphasize applied machine learning, natural language processing, and large language model-driven workflows using real-world clinical data, including electronic health record data, pathology reports, radiology reports, clinical notes, genomics, treatment data, and other institutional data sources. The Data Scientist will work semi-independently in close collaboration with clinical investigators, informatics teams, biostatisticians, and other data science stakeholders to design, build, evaluate, and refine computational pipelines. The ideal candidate will have practical prior experience developing data science workflows in Python and using modern machine learning or LLM-based tools in real projects. Job Responsibility

  • Develop, test, and maintain Python-based data pipelines for clinical research, quality improvement, and computational oncology projects.
  • Support cancer informatics projects involving natural language processing, machine learning, large language models, and structured extraction from unstructured clinical data.
  • Build workflows for processing clinical notes, pathology reports, radiology reports, treatment records, genomics reports, and other real-world healthcare data sources.
  • Implement and evaluate LLM-assisted workflows, including prompt engineering, structured output generation, model benchmarking, validation pipelines, and error analysis.
  • Assist with the development of retrieval-augmented generation workflows, vector search, embedding-based retrieval, and related approaches where appropriate.
  • Work with clinical subject matter experts to translate oncology-focused research questions into executable data science tasks.
  • Perform data cleaning, data wrangling, exploratory analysis, feature engineering, model development, and model performance evaluation.
  • Generate reproducible analyses, reports, dashboards, tables, and visualizations to communicate findings to clinical and operational stakeholders.
  • Maintain clear documentation of code, analytic decisions, model assumptions, validation methods, and project outputs.
  • Participate in model validation efforts, including comparison of computational outputs against clinician-reviewed reference standards.
  • Contribute to manuscript, abstract, grant, and presentation development through data analysis, figure generation, and methods documentation.
  • Work independently on assigned analytic tasks while communicating progress, limitations, and blockers clearly to project leadership.

Job Qualification

  • Bachelor’s Degree in Computer Science, Informatics, Statistics, Engineering, Data Science, or related field, required. Master’s Degree, preferred.
  • Minimum of two (2) years of post-graduate training or experience involving quantitative data analysis, required and working with clinical data, data science, and machine learning, preferred.
  • Working familiarity with basic medical and health information technology concepts, including standardized terminologies and ontologies and electronic health records, as well as Data Warehousing and Business Intelligence tools, required.
  • Expertise in working with SQL relational databases and statistical or general programming languages (e.g., Python, R), required.
  • Deep understanding of statistical and predictive modeling concepts, machine-learning approaches, clustering and classification techniques, and recommendation and optimization algorithms.

HIGHLY PREFERRED

  • Demonstrated prior experience building or implementing applied data science, machine learning, NLP, or LLM-based workflows. Completion of a short AI certificate, bootcamp, or introductory course alone is not sufficient for this role.
  • Strong practical experience with Python for data science, including pandas, NumPy, scikit-learn, Jupyter notebooks, and reproducible analytic workflows.
  • Prior experience applying machine learning, natural language processing, or large language models to real-world data problems.
  • Experience using off-the-shelf LLMs through APIs or enterprise platforms, including structured prompting, output parsing, evaluation, and workflow integration.
  • Experience with retrieval-augmented generation, vector databases, embeddings, semantic search, or document retrieval pipelines.
  • Experience working with clinical, biomedical, or electronic health record data.
  • Familiarity with oncology data, cancer registries, pathology reports, radiology reports, genomics reports, or clinical trial data.

Apply tot his job Apply To this Job

You might like

Senior Data Engineer – Remote, Azure & Analytics

Work from home Full-time role

Remote Data Scientist/Analyst (Entry/Junior Level)

Work from home Full-time role

Senior Data Scientist, Product Analytics (Remote, US)

Work from home Full-time role

Data Analyst/Scientist - Junior/Entry (Remote)

Work from home Full-time role

Data Scientist (medium-term contracted assignment)

Work from home Full-time role

Junior ML Engineer

Work from home Full-time role

Data Engineer, Data Platforms (Remote)

Work from home Full-time role

Remote Cloud Data Engineer

Work from home Full-time role

Data Engineering Lead - 100% Remote

Work from home Full-time role

Forward Deploy Sr. Data Engineer-VP

Work from home Full-time role

Experienced Customer Service and Communications Lead – Monitor Unit Operations

Work from home Full-time role

Experienced Live Chat Support Agent – Delivering Exceptional Customer Experiences in a Dynamic Remote Environment

Work from home Full-time role

Start Your Remote Career - Work-Life Balance | Immediate Start | No Experience Needed

Work from home Full-time role

Experienced Full Stack Data Scientist – Web & Cloud Application Development at arenaflex

Work from home Full-time role

Experienced Online Data Entry Specialist – Remote Work Opportunity with arenaflex

Work from home Full-time role

Director of Data and Business Intelligence, Need Python – Work From Home

Work from home Full-time role

Experienced Customer Service Representative – Remote Travel Support

Work from home Full-time role

Certified Peer Recovery Specialist

Work from home Full-time role

Experienced Data Entry Specialist – Entry-Level Opportunity at arenaflex

Work from home Full-time role

Senior Analyst, Protective Intelligence (Remote)

Work from home Full-time role