data engineer.
We're looking for a Data Engineer to develop and enhance application components that support ML/AI models and data ingestion pipelines. You'll lead and collaborate with a team of developers, work closely with business stakeholders and ensure that data systems and statistical models are production-ready, resilient and scalable.

function
Data & analytics / engineering
type
Full-time / contract
level
mid to senior (4+ years)
primary stack
Oracle exadata, python, hadoop, spark
what you'll do.
Three areas of ownership, from pipelines through model support and team leadership
01
Data Engineering & Pipeline Dev
→
Design, develop and maintain ETL pipelines and data ingestion processes supporting ML/AI models
→
Build and enhance Hive and DBMS-based applications for large-scale data processing
→
Develop PySpark and Spark jobs for big data transformation and analytics workloads
→
Write and optimize Oracle SQL/PLSQL stored procedures and queries on Exadata
02
ML/AI Model Support
→
Develop components that prepare and serve data for machine learning model training and inference
→
Implement and deploy statistical models using Python libraries (scikit-learn, scipy, numpy, pandas)
→
Work in Jupyter notebooks to prototype, evaluate and document model pipelines
→
Ensure models are production-ready with a focus on code resiliency and stability
03
Collaboration & Leadership
→
Lead and mentor a team of developers on data engineering best practices
→
Partner with business stakeholders to translate requirements into scalable data solutions
→
Participate in design and code peer reviews
→
Contribute to analysis of operational issues and drive resolution
→
Work in a fast-paced Agile environment with minimal supervision
what we're looking for.
the bonus list makes you stand out
Required
✓
Strong knowledge of Oracle, SQL and RDBMS systems including Oracle Exadata
✓
Hands-on experience with Hadoop, Hive, Spark and PySpark for big data workloads
✓
Python programming proficiency including scripting and object-oriented design
✓
Experience with ETL development and data pipeline architecture
✓
Familiarity with statistical modeling libraries: Jupyter, scipy, numpy, pandas, scikit-learn
✓
Working knowledge of machine learning concepts and model lifecycle
Nice to have
+
Experience automating and deploying ML models in a production environment
+
Knowledge of Autosys for job scheduling
+
Prior experience with Horizon tools: Jira, Bitbucket
+
Familiarity with Natural Language Processing (NLP) techniques including semantic search, classification, and information extraction
+
Demonstrated ability to manage senior stakeholder expectations and build trust
tech stack snapshot.
a quick scan of everything you'll touch day to day
primary skills
Oracle Exadata
secondary skill
Oracle SQL / PLSQL
tertiary skills
Hadoop
languages
Python (primary)
SQL
big data
Hadoop
Hive
PySpark
Spark
ml libraries
scipy
scikit-learn
numpy
pandas
Jupyter
ETL & pipelines
Custom ETL
DBMS-based ingestion
scheduling
Autosys
project tools
Jira
Bitbucket
Agile/Scrum