GSS Group logo

Sr. Data Engineer - PySpark (Banking)

GSS Group
11 hours ago
Full-time
On-site
Dubai, 03

JobsCloseBy Editorial Insights

Senior Data Engineer, PySpark (Banking) at GSS Group is a full-time onsite role in Dubai seeking a hands-on expert with PySpark, Python and Informatica BDM to design and run scalable ETL pipelines, data marts and enterprise-grade data platforms in a banking environment. Candidates should bring 5+ years of commercial experience, end-to-end SDLC ownership, and a track record of delivering robust data processing across structured, semi-structured, and unstructured sources, plus strong Oracle SQL and data quality practices. To apply, tailor your resume to highlight production deployments, UAT and post-production support, CI/CD and data pipeline optimization, debugging and performance tuning, cross-functional collaboration, and banking domain experience.


We are looking for an experienced Senior Data Engineer with strong expertise in PySpark, Python, and Informatica BDM development to support large-scale Data & Analytics initiatives within an enterprise banking environment. The ideal candidate should possess hands-on experience in building and maintaining robust ETL pipelines, data marts, and scalable data processing solutions across structured, semi-structured, and unstructured data sources.

The role requires strong analytical capabilities, end-to-end SDLC experience, and the ability to work closely with cross-functional teams throughout development, testing, deployment, and production support activities.

Requirements

Key Responsibilities

  • Design, develop, and maintain scalable ETL pipelines and Data Mart solutions.
  • Develop high-performance data processing solutions using PySpark and Python.
  • Perform end-to-end SDLC activities including development, UAT support, bug fixing, production deployment, and post-production support.
  • Build and optimize data transformation pipelines for large-scale enterprise data platforms.
  • Perform data analysis, code debugging, and performance tuning across PySpark and SQL-based solutions.
  • Collaborate with business, analytics, and engineering teams to understand and implement data requirements.
  • Ensure data quality, integrity, scalability, and reliability across data pipelines.
  • Work with structured, semi-structured, and unstructured datasets in enterprise environments.
  • Participate in code reviews and implement software engineering best practices.
  • Support CI/CD implementation and data pipeline deployment activities.
  • Troubleshoot production issues and implement effective resolutions.
  • Contribute to technical documentation and knowledge-sharing initiatives.

Required Technical Skills

Programming & Big Data Technologies

  • Python
  • PySpark
  • Apache Spark
  • Informatica BDM (Big Data Management)
  • Hadoop
  • MapReduce
  • Hive
  • Pandas

Data Engineering & ETL

  • ETL Pipeline Development
  • Data Mart Development
  • Data Warehousing Concepts
  • Data Transformation & Processing
  • Data Pipeline Optimization

Databases & Query Languages

  • Oracle SQL
  • SQL
  • NoSQL Databases
  • Strong analytical and query-writing skills

Tools & Platforms

  • Jupyter Notebook
  • Git / Version Control
  • CI/CD Pipelines
  • Testing & Validation Frameworks

Required Experience

  • Minimum 5+ years of commercial experience in Data Engineering or Data Analytics projects.
  • Strong hands-on experience in PySpark and Python-based ETL development.
  • Experience building enterprise-grade ETL pipelines and Data Mart solutions.
  • Strong experience in Informatica BDM development.
  • Experience handling end-to-end SDLC activities including development, UAT, production deployment, and post-production support.
  • Strong expertise in Oracle SQL and data analysis.
  • Hands-on experience debugging PySpark code and optimizing data processing workflows.
  • Experience working with production-grade data pipelines and large datasets.
  • Strong understanding of software engineering principles and coding best practices.
  • Experience working with Agile delivery environments.

Preferred Domain Experience

  • Banking
  • Financial Services
  • Digital Products
  • Data & Analytics Platforms

Nice to Have

  • Experience working with enterprise data lake and big data ecosystems.
  • Exposure to cloud-based data platforms.
  • Experience working with CI/CD and automated data pipeline deployments.
  • Knowledge of modern data engineering best practices and scalable architectures.

Daily Tech Stack

The selected candidate will work extensively with:

  • Python
  • PySpark
  • Informatica BDM
  • Apache Spark
  • Jupyter Notebook
  • Oracle SQL
  • SQL & NoSQL Databases
  • Hadoop Ecosystem (Hive, MapReduce)
  • CI/CD Tools
  • ETL & Data Warehousing Technologies

Functional Competencies

  • Strong problem-solving and analytical skills.
  • Excellent debugging and troubleshooting capabilities.
  • Ability to work independently in a fast-paced Agile environment.
  • Strong ownership mindset and attention to detail.
  • Effective stakeholder communication and collaboration skills.
  • Ability to manage multiple priorities and production support activities.