Skip to content

server : fix endpoint checks - #10135

Merged
ggerganov merged 1 commit into
masterfrom
gg/server-fix-endpoints
Nov 2, 2024
Merged

ggerganov merged 1 commit into
masterfrom
gg/server-fix-endpoints

Conversation

@ggerganov

Copy link
Copy Markdown
Member

ref #3815 (comment)

I think 0d6f6a7 messed up the endpoint checks. Per the readme, the --embeddings flag should restrict to just /embeddings endpoint, while --reranking should enable the /rerank endpoint.

@ngxson

ngxson commented Nov 2, 2024

Copy link
Copy Markdown
Collaborator

Hmm yeah I think I misunderstood --embedding. I thought that it means "enable embd && disable completion"

So just to confirm, the server does support having some slots running embd and some slots running completion at the same time, right?

I'm asking this because I can't find llama_set_causal_attn anywhere in the server.cpp code

@ggerganov

Copy link
Copy Markdown
Member Author

So just to confirm, the server does support having some slots running embd and some slots running completion at the same time, right?

I'm asking this because I can't find llama_set_causal_attn anywhere in the server.cpp code

I'm not really sure what is the state of this functionality, and AFAIK most people use the "embedding + completion" mode just for testing purposes (i.e. avoid starting 2 separate instances of llama-server). Technically, for getting the embeddings from a LLaMA model for example, you don't need to call llama_set_causal_attn(false). Just llama_set_embeddings(true) which we already do. There are models like GritLM which would require correct calls to llama_set_causal_attn and this is not supported by llama-server atm.

@ngxson ngxson left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK thanks for the explanation. That sounds good.

@ggerganov
ggerganov merged commit 4595041 into master Nov 2, 2024
@ggerganov
ggerganov deleted the gg/server-fix-endpoints branch November 2, 2024 16:34
@ggerganov ggerganov mentioned this pull request Nov 5, 2024
4 tasks done
arthw pushed a commit to arthw/llama.cpp that referenced this pull request Nov 15, 2024
arthw pushed a commit to arthw/llama.cpp that referenced this pull request Nov 18, 2024
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
ljubomirj pushed a commit to ljubomirj/llama.cpp that referenced this pull request May 6, 2026
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
AlexiAlp pushed a commit to minghaop/llama.cpp that referenced this pull request Jun 2, 2026
AlexiAlp pushed a commit to minghaop/llama.cpp that referenced this pull request Jun 2, 2026
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants