Summary from listing
The role is to design and build an autonomous AI/ML framework that predicts failures in storage, network, server delivery, and VMware infrastructure by ingesting data from Elastic Search, Splunk, OSWatcher, Zabbix, Snow Inc, TFS Defects, and other event-log sources. The framework will preprocess and store data in an ELK cluster, train and continuously improve failure-prediction models, and provide remediation inputs to StackStorm. The work follows CI/CD and uses Python, Keras, TensorFlow, Flask, MongoDB, Elastic Search, Splunk, Hadoop, and PostgreSQL.
