Information Technology_USA - USA_Engineer
Engineering
2 Candidate Submittal Slots, New High Level PolicyBill Rate - , some flexibility for extremely well qualified candidatesMSP Owner: Rob FintonLocation: Seattle, Washington, 98109 - 100% OnsiteDuration: 6 monthsGBaMS ReqID: 10928911Competencies: 6+ years experience requiredDigital : Amazon Web Service(AWS) Cloud ComputingAdvanced Java ConceptsMicrosoft SQL Server 2019Java Performance Tools (Jprobe, Jensor, OptimizeIT)Must Have Technical/Functional Skills• AWS EMR Cluster Operations • Spark & YARN Tuning • Memory Optimization & Capacity Analysis • EC2 Right-Sizing & Cost Optimization • Monitoring & Troubleshooting (CloudWatch/Logs)Roles & Responsibilities:• Analyze EMR cluster metrics, Spark application telemetry, and YARN resource utilization to identify over-provisioned memory allocations, underutilized executors, and suboptimal cluster configurations that contribute to excessive DRAM consumption.• Recommend and implement cluster-level optimizations, including instance family right-sizing (e.g., migrating from memory-optimized R-type to compute-optimized C-type instances), node count adjustments, EBS volume configurations, and spot/on-demand fleet composition changes.• Tune Spark runtime configurations at the cluster level, including executor memory/core ratios, YARN container sizing, dynamic resource allocation settings, memory overhead parameters, and shuffle service configurations, to achieve optimal memory utilization without impacting job SLAs.• Perform custom operations and iterative experiments using Amazon internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results.• Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders.• Monitor service health metrics and troubleshoot operational issues during and after optimization activities, ensuring zero degradation to job completion times, data processing throughput, and downstream SLAs.• Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture cluster optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale.• Extract scalable learnings from optimization engagements and develop programmatic frameworks that enable the initiative to scale across hundreds of EMR clusters, including training and enabling other vendor engineers to execute optimization playbooks.
