Posted 22 July, 2026
Devops Engineer
The Mother Ship
San Antonio, TX, US
Full Time
Job Description
Location: San Antonio
\nDevops Engineer
\nTools & Technologies:
\n- \n
- Apache Kafka (Self-managed or MSK) \n
- AWS managed Apache Flink \n
- Amazon EC2, S3, RDS, and VPC \n
- Terraform/CloudFormation \n
- Docker, Kubernetes (EKS) \n
- Elk, CloudWatch \n
- Python, Bash \n
Skills and Expertise:
\n- \n
- AWS Managed Services: \n
- \n
- Proficiency in AWS services such as Amazon MSK (Managed Streaming for Kafka), Amazon Kinesis, AWS Lambda, Amazon S3, Amazon EC2, Amazon RDS, Amazon VPC, and AWS IAM. \n
- Ability to manage infrastructure as code with AWS CloudFormation or Terraform. \n
- \n
- Apache Flink: \n
- \n
- Understanding of Apache Flink for real-time stream processing and batch data processing. \n
- Familiarity with Flinks integration with Kafka, or other messaging services. \n
- Experience in managing Flink clusters on AWS (using EC2, EKS, or managed services). \n
- \n
- Kafka Broker (Apache Kafka): \n
- \n
- Deep knowledge of Kafka architecture, including brokers, topics, partitions, producers, consumers, and zookeeper. \n
- Proficiency with Kafka management, monitoring, scaling, and optimization. \n
- Hands-on experience with Amazon MSK (Managed Streaming for Kafka) or self-managed Kafka clusters on EC2. \n
- \n
- DevOps & Automation: \n
- \n
- Strong experience in automating deployments and infrastructure provisioning. \n
- Familiarity with CI/CD pipelines using tools like Jenkins, GitLab, GitHub Actions, CircleCI, etc. \n
- Experience with Docker and Kubernetes, especially for containerizing and orchestrating applications in cloud environments. \n
- \n
- Programming & Scripting: \n
- \n
- Strong scripting skills in Python, Bash, or Go for automation tasks. \n
- Ability to write and maintain code for integrating data pipelines with Kafka, Flink, and other data sources. \n
- \n
- Monitoring & Performance Tuning: \n
- \n
- Knowledge of CloudWatch, Prometheus, Grafana, or similar monitoring tools to observe Kafka, Flink, and AWS service health. \n
- Expertise in optimizing real-time data pipelines for scalability, fault tolerance, and performance. \n
Responsibilities:
\n- \n
- Infrastructure Design & Implementation: \n
- \n
- Design and deploy scalable and fault-tolerant real-time data processing pipelines using Apache Flink and Kafka on AWS. \n
- Build highly available, resilient infrastructure for data streaming, including Kafka brokers and Flink clusters. \n
- \n
- Platform Management: \n
- \n
- Manage and optimize the performance and scaling of Kafka clusters (using MSK or self-managed). \n
- Configure, monitor, and troubleshoot Flink jobs on AWS infrastructure. \n
- Oversee the deployment of data processing workloads, ensuring low-latency, high-throughput processing. \n
- \n
- Automation & CI/CD: \n
- \n
- Automate infrastructure provisioning, deployment, and monitoring using Terraform, CloudFormation, or other tools. \n
- Integrate new applications and services into CI/CD pipelines for real-time processing. \n
- \n
- Collaboration with Data Engineering Teams: \n
- \n
- Work closely with Data Engineers, Data Scientists, and DevOps teams to ensure smooth integration of data systems and services. \n
- Ensure the data platforms scalability and performance meet the needs of real-time applications. \n
- \n
- Security and Compliance: \n
- \n
- Implement proper security mechanisms for Kafka and Flink clusters (e.g., encryption, access control, VPC configurations). \n
- Ensure compliance with organizational and regulatory standards, such as GDPR or HIPAA, where necessary. \n
- \n
- Optimization & Troubleshooting: \n
- \n
- Optimize Kafka and Flink deployments for performance, latency, and resource utilization. \n
- Troubleshoot issues related to Kafka message delivery, Flink job failures, or AWS service outages. \n
