[DRAFT] Model serve multiple backends - #1386
Conversation
rewrite/copy ensure_serve logic to work with a model served with vllm for `ilab chat` Signed-off-by: Ali Maredia <amaredia@redhat.com>
Signed-off-by: Ali Maredia <amaredia@redhat.com>
Signed-off-by: Ali Maredia <amaredia@redhat.com>
Signed-off-by: Ali Maredia <amaredia@redhat.com>
|
The vLLM package on PyPI is a CUDA build. How do you plan to support CPU-only, ROCm, and the vLLM fork for Intel Gaudi? |
|
@tiran I understand your concerns and your comments in #1276 were insightful and will be applied/taken into consideration. The requirements.txt in this file was able to work in my development environment, and is not intended to be merged. This PR is purely a draft, @leseb will be taking over this integrating vllm into ilab. I just wanted a place where many people can get eyes on the code and discuss the feature. |
|
This pull request has merge conflicts that must be resolved before it can be |
|
Closing in favor of #1442. Thanks! |
Checklist:
conventional commits.