> For the complete documentation index, see [llms.txt](https://sliu583.gitbook.io/blog/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://sliu583.gitbook.io/blog/specific-work/shivarams-group/group-papers/fluid-resource-aware-hyperparameter-tuning-engine.md).

# Fluid: Resource-aware Hyperparameter Tuning Engine

https\://proceedings.mlsys.org/paper/2021/file/9f61408e3afb633e50cdf1b20de6f466-Paper.pdf

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsLdilqvzEu-pJHjfJ%2Fimage.png?alt=media\&token=6498b159-1f3d-4b5e-94c9-62d9b78c2eac)

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsLj-hsiz8S3S0ACvx%2Fimage.png?alt=media\&token=d2fce56a-01e6-4db6-82e3-83c7f8e9afbf)

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsLlyDl7ZygPjcQ-R2%2Fimage.png?alt=media\&token=7889bd2c-3ecf-4730-9316-353d85f17f2b)

* Successive Halving&#x20;
  * Workers underutilized, only one job to run&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsMCVAuMmWpntdBsPG%2Fimage.png?alt=media\&token=cdcc0ccf-7fb4-4b75-8353-cfe0157ef245)

* More efficient, but there're also some problems&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsMOgvx4UN4Fhujweo%2Fimage.png?alt=media\&token=4ea4c3fe-820a-4237-949f-4b4bf3b26d7c)

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsMZSMbm42jQLMnt76%2Fimage.png?alt=media\&token=5ea4735c-c975-4e24-86a4-437f83438c22)

* Goal: Utilized & Useful work?&#x20;
* Resource-aware hyperparameter tuning&#x20;
  * Previous work: Hypersched (resource management), but more specialized on the algorithm itself&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsNHk2H1HUnapDFSQK%2Fimage.png?alt=media\&token=8787565e-3393-469d-a3b2-a1c1f4e8a6cf)

* Intra-GPU sharing (pack jobs in single GPU)
  * Current schedule: FIFO queue to manage to jobs&#x20;

### Design and Algorithms&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsNyD7d1-OoFVq4qCa%2Fimage.png?alt=media\&token=26364439-6cfd-419e-925b-9df07c7d7406)

* Intra: packing several jobs on single GPU&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsOK6nR9yBv7VlXRKE%2Fimage.png?alt=media\&token=69ad4ed9-4168-4a11-8e36-275395752f00)

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsOzRwLDey2u7ZBkat%2Fimage.png?alt=media\&token=532d103f-5a5e-43e5-9fc0-07c285a9b947)

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsPLLC0usKwHa2bk5m%2Fimage.png?alt=media\&token=9ed29735-57cc-4e58-a1c6-c675bfa74c84)

* Very helpful when people read the paper (very helpful in communicating the ideas)&#x20;
* Overhead of packing?&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsPiBoqoxTFdPrtLW0%2Fimage.png?alt=media\&token=bfbf305b-413c-4184-9de2-a2bc7756377b)

* Packing overhead of doing the placement&#x20;
  * Limit the number of packing trails&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsPxs4Wz1IQFO-c61H%2Fimage.png?alt=media\&token=bc5ef738-6700-4248-bdb4-8f36d9b9aab4)

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsREk6jbPfkpdew3tK%2Fimage.png?alt=media\&token=fee59e80-9d33-4171-97ff-b872c3282163)

* Model fits on the single worker&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsRjkJfCOwlKwSvObA%2Fimage.png?alt=media\&token=56f72ef7-14d9-4c9f-9fa5-91b94ec57e97)

* Makespan of all the trials&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsRoFE3Wco18ldEkQ6%2Fimage.png?alt=media\&token=9429368c-b144-4108-9bc3-d1b97858e4f0)

* Improvement more prominent when applying for asynchronous version&#x20;
* Intuition:&#x20;
  * If the runtime is very scaled&#x20;
* Parameters they are tuning
  * Learning rate, dropout rate&#x20;
  * Number of layers&#x20;
  * Batch size&#x20;
* Failures&#x20;
  * Know reasonable ranges&#x20;

![](https://2097630930-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MVORxAomcgtzVVUqmws%2F-MZsJwsO-aGwSAMubWap%2F-MZsTJ5pXbzo7uUn1ah4%2Fimage.png?alt=media\&token=9c84372c-302a-4b2a-b7fe-73466eed932a)

* Multiple hyperparameter jobs&#x20;
  * Multiple trial groups&#x20;
  * But extra parallelism there?&#x20;
  * Different things in the same trial group&#x20;
* Variability&#x20;
  * Some&#x20;
  * More?&#x20;
* Space sharing & Parallelism&#x20;
  * Automatically parallelism, don't need anything from hyperparameter&#x20;
  * Queue problem, but in tight space&#x20;
