Data Engineer I

Dun & Bradstreet | Hyderabad, IN | hybrid

Apply for the Data Engineer I role at Dun & Bradstreet in Hyderabad, India. Full-time hybrid opportunity focused on Python, SQL, ETL/ELT, data pipelines, Playwright, Selenium, BigQuery, AWS/GCP, LangChain, RAG, vector databases and AI-enabled data workflows.

Job description

Dun & Bradstreet is hiring a Data Engineer I for a full-time, hybrid role in Hyderabad, India. The engineer will work as part of an Agile team supporting Research and Managed Services in a combined technical and operational capacity. The role involves writing and maintaining code for data tools, automations, ingestion and transformation workflows, evaluating data sources for AI-enabled and traditional research workflows, and helping keep production data operations reliable. Responsibilities Write, review, test, and maintain SQL and Python code for data tools, automations, ingestion workflows, and data transformations. Automate manual processes to improve efficiency, quality, accuracy, and throughput. Develop coding standards and contribute to peer code reviews. Build and maintain web-scraping solutions, API integrations, and reusable data-processing components. Support scalable ETL/ELT pipelines for structured and unstructured data. Identify and evaluate internal and external data sources for AI-enabled and traditional research workflows. Profile and validate sources for relevance, authority, freshness, completeness, accessibility, reliability, legal constraints, privacy, security, and technical compatibility. Document metadata, source decisions, lineage, ownership, refresh expectations, limitations, and approved use cases. Implement and support AI-enabled workflows using LangChain or equivalent orchestration frameworks, LLMs, embeddings, RAG, vector databases, and prompt-engineering techniques where applicable. Monitor data-source and AI-workflow performance and recommend remediation or replacement when quality drops. Support day-to-day operations including monitoring, exception handling, data maintenance, and issue resolution. Perform database administration, performance tuning, and data-quality maintenance. Investigate production incidents, pipeline failures, data-quality issues, and operational exceptions. Maintain technical documentation including data dictionaries, data-flow diagrams, mappings, runbooks, and lineage records. Collaborate with teams across Technology, Data & Analytics, Research Services, Managed Services, Product, and Data Governance.

Responsibilities

  • Write, review, test, and maintain SQL and Python code for data tools, automations, ingestion workflows, and data transformations.
  • Automate manual processes to improve efficiency, quality, accuracy, and throughput.
  • Develop coding standards and contribute to peer code reviews.
  • Build and maintain web-scraping solutions, API integrations, and reusable data-processing components.
  • Support scalable ETL/ELT pipelines for structured and unstructured data.
  • Identify and evaluate internal and external data sources for AI-enabled and traditional research workflows.
  • Profile and validate sources for relevance, authority, freshness, completeness, accessibility, reliability, legal constraints, privacy, security, and technical compatibility.
  • Document metadata, source decisions, lineage, ownership, refresh expectations, limitations, and approved use cases.
  • Implement and support AI-enabled workflows using LangChain or equivalent orchestration frameworks, LLMs, embeddings, RAG, vector databases, and prompt-engineering techniques where applicable.
  • Monitor data-source and AI-workflow performance and recommend remediation or replacement when quality drops.
  • Support day-to-day operations including monitoring, exception handling, data maintenance, and issue resolution.
  • Perform database administration, performance tuning, and data-quality maintenance.
  • Investigate production incidents, pipeline failures, data-quality issues, and operational exceptions.
  • Maintain technical documentation including data dictionaries, data-flow diagrams, mappings, runbooks, and lineage records.
  • Collaborate with teams across Technology, Data & Analytics, Research Services, Managed Services, Product, and Data Governance.

Requirements

  • Strong SQL and Python skills, with demonstrated ability to write and maintain production code as part of daily work.
  • Experience with Playwright, Selenium, and other web-data-collection techniques.
  • Experience developing and supporting data-ingestion, transformation, and ETL/ELT workflows.
  • Ability to collect and interpret data from multiple sources including web scraping, GCS/S3, delimited files, XML, JSON, and PDF.
  • Working knowledge of data systems and databases used to maintain data pipelines.
  • Experience with Power BI, Tableau, or other dashboard tools.
  • Experience managing stakeholders and project plans.
  • Proficiency with Microsoft Office.
  • BigQuery experience and knowledge of AWS and/or GCP.
  • Hands-on experience implementing AI solutions with LangChain or an equivalent framework.
  • Exposure to LLMs, prompt engineering, RAG, embeddings, vector databases, AI agents, or graph databases.
  • Knowledge of Data Operations methodologies, data-management approaches, ServiceNow, and/or Jira.
  • Experience with NoSQL, SQL Server administration, R, web technologies, or multi-source data mapping is relevant.
  • Ability and willingness to learn and adopt new technologies.

Skills

  • Data Engineering
  • Python
  • SQL
  • ETL
  • ELT
  • Data Pipelines
  • Web Scraping
  • Playwright
  • Selenium
  • API Integrations
  • BigQuery
  • AWS
  • GCP
  • Power BI
  • Tableau
  • LangChain
  • Large Language Models
  • RAG
  • Embeddings
  • Vector Databases
  • AI Agents
  • NoSQL
  • SQL Server
  • Data Governance
  • Data Quality
  • ServiceNow
  • Jira