First think I want to say, I had hard time connecting to Ollama because it expects the request for chat completions to be sent to http://localhost:11434/v1/ endpoint. I had it configured to http://localhost:11434 in Open-Interface and had to debug+google a lot to figure out that I need to configure it to http://localhost:11434/v1/ (so maybe adding info about this to Readme would be beneficial?). After that, I was able to see in the Ollama server output that Open-Interface did sent the correct request. However, I only tried it with the moondream model, and it failed on the step of converting the response from the LLM to json instructions. I guess the problem is that I use just too "dumb" model.
So my question is if there is someone who managed to make the Open-Interface work with some small local multimodal model? The model does not need to be compatible with Ollama, I am open to installing other inference software, but I would like it to fit to my Nvidia A500 RTX 4GB graphics card. I'm also ok with larger models that will use also my CPU, but the fitting on GPU would be the best.
My OS is Ubuntu 24.04. I wonder if Open-Interface is able to open the chrome/firefox if it does not see it on the screen, as my dock is only visible on hower. But I guess it can click "Win" and then type chrome, but I do not know, I did not yet dig deep enough into the source code.
So let the discussion begin. I also noticed the Discussions tab here on the github, but I feel like it is a dead space. I never used that, so I am putting my topic for discussion here, hope you don't mind.
First think I want to say, I had hard time connecting to Ollama because it expects the request for chat completions to be sent to
http://localhost:11434/v1/endpoint. I had it configured tohttp://localhost:11434in Open-Interface and had to debug+google a lot to figure out that I need to configure it tohttp://localhost:11434/v1/(so maybe adding info about this to Readme would be beneficial?). After that, I was able to see in the Ollama server output that Open-Interface did sent the correct request. However, I only tried it with the moondream model, and it failed on the step of converting the response from the LLM to json instructions. I guess the problem is that I use just too "dumb" model.So my question is if there is someone who managed to make the Open-Interface work with some small local multimodal model? The model does not need to be compatible with Ollama, I am open to installing other inference software, but I would like it to fit to my
Nvidia A500 RTX 4GBgraphics card. I'm also ok with larger models that will use also my CPU, but the fitting on GPU would be the best.My OS is
Ubuntu 24.04. I wonder if Open-Interface is able to open the chrome/firefox if it does not see it on the screen, as my dock is only visible on hower. But I guess it can click "Win" and then type chrome, but I do not know, I did not yet dig deep enough into the source code.So let the discussion begin. I also noticed the
Discussionstab here on the github, but I feel like it is a dead space. I never used that, so I am putting my topic for discussion here, hope you don't mind.