Data & Knowledge Engineer
Job Summary
Job Description & Summary
Provide trusted contextual and well-governed enterprise data and knowledge services that ground agentic workflows and improve their reliability.
Design and build ingestion transformation and serving pipelines for structured and unstructured data.
Create retrieval indexes metadata models semantic layers knowledge graphs or data products as appropriate.
Implement chunking enrichment lineage quality and access-control patterns.
Optimize retrieval quality freshness latency and cost with the AI engineering team.
Integrate cloud and on-premises data sources for hybrid solutions.
Support evaluation datasets monitoring data and traceability requirements.
4 years in data engineering analytics engineering information retrieval or knowledge platforms.
Strong SQL and Python skills and experience with data pipelines APIs and data modeling.
Practical knowledge of vector search embeddings metadata document processing and retrieval evaluation.
Experience with enterprise security data quality and hybrid data integration.
Strong SQL and Python capability with practical experience in Spark and data engineering platforms such as Microsoft Fabric Azure Data Factory Databricks Snowflake or equivalent.
Hands-on experience processing structured and unstructured content including parsing OCR chunking enrichment metadata extraction lineage and incremental indexing.
Experience with vector and hybrid search technologies such as Azure AI Search PostgreSQL with pgvector Elasticsearch Pinecone Weaviate Milvus or equivalent.
Understanding of embedding selection semantic and lexical retrieval metadata filtering reranking query transformation evaluation datasets and retrieval quality metrics.
Experience with graph and knowledge technologies such as Neo4j RDF or property graphs ontologies entity resolution and GraphRAG patterns is desirable.
Ability to implement secure hybrid data access row or document-level permissions data masking and traceable ingestion from cloud and on-premises repositories.
Data freshness quality and availability
Retrieval relevance and traceability
Speed of onboarding new knowledge sources
Pipeline reliability and performance
Compliance with data-access requirements
Other members of the AI Transformation & Agentic Systems Practice
PwC sector functional cloud cyber risk Responsible AI and change specialists
Client business owners product owners technology teams and operational users
Technology alliance and implementation partners where relevant
Support proposals client workshops and market development appropriate to seniority.
Contribute reusable methods patterns code assets and lessons learned.
Coach colleagues and participate in the capabilitys continuous learning agenda.
Uphold PwC quality independence confidentiality and risk-management requirements.
#LI-BS1 #LI-Hybrid
Required Experience:
IC
About Company
At PwC, our purpose is to build trust in society and solve important problems. We’re a network of firms in 155 countries with over 284,000 people who are committed to delivering quality in assurance, advisory and tax services. Find out more and tell us what matters to you by vis ... View more