About the team: We are the Big Data Infrastructure team at Binance, responsible for the stability and cost efficiency of the company's core data platform. Our systems cover Spark/YARN job monitoring and governance, S3 storage cost analysis, data lineage and quality management, and automated operations tooling.
Responsibilities
Spark / YARN Job Optimization
Analyze job resource usage and identify inefficiencies (CPU idle, memory waste, timeouts, small file problems, etc.)
Build job profiling systems: job classification, resource baseline modeling, historical trend analysis
Produce optimization reports and drive business owners to implement improvements
AI-Driven Automation & Tooling
Use AI coding tools (Cursor / Copilot / Claude, etc.) fluently to accelerate development and tool delivery
Participate in EMR / S3 cost optimization projects: analyze high-cost jobs and storage, identify waste, and drive owners to take action
Develop automated ops scripts: scheduled health checks, anomaly alerting, data governance policy auto-deployment
Help build internal SaaS tooling to systematize and productize repetitive manual operations
Monitoring & Alerting
Maintain and improve existing monitoring systems (Prometheus metrics, alerts, log analysis)
Participate in Flink / Spark job health monitoring development
Assist with capacity alerting and governance for disk, S3, and other storage layers
What We're Looking For
Currently enrolled in a Bachelor's or Master's program in Computer Science, Software Engineering, or a related field
Proficient in Python — able to write clean, maintainable scripts and tools independently
**Must be fluent in AI coding tools (Cursor / Copilot / ChatGPT / Claude, etc.) for development and troubleshooting** — this is a hard requirement; please think carefully before applying if you don't use AI tools regularly
Comfortable with Linux: running programs on servers, reading logs, debugging issues
Bonus Points (one or two is plenty)
Have used Spark / Hive / Flink, even just in coursework
Familiar with AWS basics (S3, EMR, Athena)
Have written Shell scripts or set up crontab scheduled tasks
Have used Kafka, Prometheus, or any message queue / monitoring system
Have completed a full project using AI tools and can clearly explain how and where you used them
Any real project experience: open source contributions, competitions, lab projects, etc.
Bilingual English/Mandarin is an added advantage to be able to coordinate with overseas partners and stakeholders.