Description
Anyscale is hiring a Ray Data Engineer to build, optimize, and scale the Ray Datasets library and data-processing capabilities. The role focuses on large-scale performance, integration with machine learning training and data sources, stability and stress testing, streaming workloads such as Beam on Ray, and differentiating data operations in Anyscale’s hosted Ray service. Responsibilities include developing open-source software, improving Ray core and Datasets architecture, enhancing the testing process, and communicating technical work through talks, tutorials, and blog posts. The role requires at least five years of relevant experience, a background in algorithms, data structures, and system design, experience with scalable and fault-tolerant distributed systems, and experience with data processing and database internals including Spark or Dask.
