Description
Description
We are seeking a Data Operations Engineer to design, build, and maintain real-time and near-realtime data ingestion pipelines supporting a mission-critical system and moves data across security domains using cross domain solutions (CDS) / guards. You will be responsible for the reliable flow of streaming data from a wide range of sources into our data platform, ensuring data quality, observability, and scalability.
You'll partner closely with data engineers, platform engineers, analytics teams, and security/governance personnel to deliver trustworthy, low-latency data that supports operational decisions — while ensuring compliance with cross-domain and data classification requirements. This position is on-site in Honolulu, HI.
Key Responsibilities
• Aid the team in delivering continual data feeds to users and monitoring the health of data quality and overall data ingest.
• Design, build, and maintain resilient real-time/near-realtime ingestion pipelines using Apache NiFi and REST APIs, with appropriate backpressure, prioritization, retries, and error-handling strategies.
• Configure and monitor data flows moving across security domains via cross domain guards/solutions, ensuring data integrity and compliance with transfer policies.
• Parse, transform, validate, and route structured and unstructured data in a variety of formats, including Excel, CSV, JSON, and XML.
• Employ data manipulation and visualization tools (e.g., Grafana, Prometheus) to effectively convey pipeline status, data quality, and historical trends to leadership, users, and the data team.
• Collaborate with platform, software, and other data engineers to (re)configure and continuously improve data ingestion pipeline reliability.
• Develop and maintain software/scripts to automate monitoring of real-time feeds and alert on timeliness, volume, lineage, and distribution data issues.
• Translate learnings from historical pipeline data into actionable steps to improve data ingest reliability and performance.
• Partner with security and governance teams to enforce encryption, authentication, authorization, and data classification requirements across pipelines and cross-domain transfers.
• Write and maintain scripts (Python, Bash, or similar) to automate data processing, validation, and monitoring tasks.
• Document data flows, system configurations, and standard operating procedures.
Qualifications
TYPICAL EDUCATION AND EXPERIENCE: Bachelors and five (5) years or more experience; Masters and three (3) years or more experience; PhD and 0 years related experience.
• U.S. Citizenship and an active TS/SCI clearance.
• Bachelor of Science degree in Computer Science, Mathematics, Electrical Engineering, Physics, Information Systems, Information Technology, or related field.
• 3+ years of experience in data engineering, data operations, or DevOps roles supporting production data pipelines.
• Proficiency in Python, Bash, or similar scripting languages commonly used in data science/data analytics applications.
• Working knowledge of the Linux (RedHat) command line; familiarity with Windows environments.
• Solid understanding of data engineering fundamentals: data pipelines, streaming architectures, ETL/ELT concepts, and data quality principles.
• Familiarity with JSON, XML, CSV, and Excel data formats, parsing, and transformation.
Preferred Qualifications (Nice to Have)
• Experience with streaming/messaging platforms and tools: Kafka, JMS, Apache Flink/Spark Streaming.
• Experience with monitoring/observability tools: Grafana, Prometheus, Elasticsearch.
• Experience with data platforms such as Snowflake.
• Familiarity with cross domain solutions/guards (e.g., data diode concepts, transfer validation, content filtering).
• Experience enforcing encryption, authentication/authorization, and data classification policies in partnership with security/governance teams.
• Experience with version control tools (e.g., Git) and Agile/Scrum practices.
• Prior experience in a government, defense, or intelligence community environment.
Apply on company website