The right talent can transform your business—and we make that happen. At Collabera, we go beyond staffing to deliver strategic workforce solutions that drive growth, innovation, and agility. With deep industry expertise, a global talent network, and a people-first approach, we connect you with professionals who don’t just fit the role but elevate your business. Partner with us and build a workforce that powers success.
Data Engineer
Remote: Houston , Texas, US span>
Salary Range: 135000.00 - 165000.00 | Per Annum
Job Code: 371658
End Date: 2026-10-15
Days Left: 26 days, 8 hours left
Location: US – Remote (EST/CST)
Employment: Full-Time opportunity
Salary Range: $135k/ann.-$165k/ann.
Key Responsibilities:
- Design, build, and maintain scalable, self-healing data pipelines using Databricks, Python, PySpark, and SQL.
- Develop data pipelines across Bronze, Silver, and Gold/Certified Gold layers, ensuring data quality and reliability.
- Ingest, transform, and serve data from dozens of source systems, including enterprise applications, financial systems, IoT, web/mobile analytics, and third-party platforms.
- Build error recovery and quarantine workflows to isolate failed records while allowing valid data to continue through the pipeline.
- Implement schema validation, data quality checks, anomaly detection, automated testing, regression testing, and data observability.
- Develop and maintain data models with proper data grain, keys, referential integrity, lineage, and governance.
- Build infrastructure that supports AI/ML workloads, including feature stores, embedding pipelines, vector search, and real-time serving layers.
- Build and maintain MCP server integrations that expose enterprise data to LLM-powered tools and AI agents.
- Develop integrations that allow AI applications to connect to, query, and retrieve data through MCP servers.
- Design and implement RAG (Retrieval-Augmented Generation) architectures and integrate LLM-powered solutions with enterprise data.
- Work with vector databases such as Pinecone, Weaviate, or similar technologies.
- Support AI model training, evaluation, deployment, monitoring, and productionization in partnership with Data Science and Product teams.
- Evaluate AI-powered data engineering and data quality tools, including solutions for automated schema detection, cataloging, completeness checks, and data validation.
- Use AI-assisted testing and evaluation approaches to validate data and AI outputs against business requirements.
- Develop APIs and data interfaces that enable AI products and internal applications to query and interact with data in real time.
- Implement data governance practices covering access controls, PII handling, data classification, compliance, and appropriate AI data usage.
- Build monitoring, alerting, SLA tracking, and data freshness capabilities into data platforms.
- Document data models, pipeline architectures, AI integrations, and reusable engineering patterns.
- 5+ years of experience in Data Engineering, working with data from multiple enterprise source systems.
- Strong hands-on experience with Databricks.
- Strong Python and PySpark development experience.
- Strong SQL skills.
- 2–3+ years of hands-on experience with MCP Servers / Model Context Protocol.
- Hands-on experience building and implementing RAG models/architectures.
- Experience connecting to, querying, and integrating data through MCP servers.
- Strong experience building self-healing or fault-tolerant data pipelines and error recovery workflows.
- Experience with schema validation and data quality frameworks.
- Strong understanding of Bronze/Silver/Gold data architecture.
- Experience with LLM integration patterns, AI agents, tool-use frameworks, and AI-enabled data solutions.
- Experience with vector databases such as Pinecone, Weaviate, or equivalent.
- Experience building data pipelines and infrastructure suitable for AI/ML workloads.
- Understanding of data governance, lineage, monitoring, observability, and data quality.
This is a direct hire opportunity. The selected candidate will be employed directly by our client. All compensation and benefits, including but not limited to medical insurance, retirement plans, paid time off, and other perks, will be provided by the client in accordance with their internal policies and subject to applicable laws and eligibility requirements.
Job Requirement
- MCP
- RAG
- Pyspark
Reach Out to a Recruiter
- Recruiter
- Phone
- Christin Mathew
- christin.mathew@collabera.com