Fast inference platform for open-source LLMs with sub-100ms latency and fine-tuning capabilities.
We couldn't find that tool. It may have been removed or the link may be incorrect.