Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
abdullah-cod9
/
llama-cpp-python
Public
forked from
JamePeng/llama-cpp-python
Notifications
You must be signed in to change notification settings
Fork
0
Star
1
Code
Pull requests
0
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Pull requests
Actions
Projects
Security and quality
Insights
Commits
Branch selector
main
User selector
All users
All time
Commit history
Commits on Apr 6, 2026
fix(chat): support tool_choice in Qwen3.5-VL chat handler
Show description for cd9388a
abdullah-cod9
committed
cd9388a
View commit details
Copy full SHA for cd9388a
Browse repository at this point
Bump version to 0.3.35
Show description for e1ade17
JamePeng
committed
e1ade17
View commit details
Copy full SHA for e1ade17
Browse repository at this point
Update README.md
JamePeng
committed
232092e
View commit details
Copy full SHA for 232092e
Browse repository at this point
fix: expand stop sequences for Gemma4ChatHandler
Show description for d4fef6a
JamePeng
committed
d4fef6a
View commit details
Copy full SHA for d4fef6a
Browse repository at this point
Relocate LFM25VLChatHandler, add comment details, and update README.md
JamePeng
committed
d7478de
View commit details
Copy full SHA for d7478de
Browse repository at this point
Merge pull request #105 from TAO71-AI/lfm2.5-vl
Show description for 9e749f9
JamePeng
authored
9e749f9
View commit details
Copy full SHA for 9e749f9
Browse repository at this point
Update Submodule vendor/llama.cpp b863507..58190cc
JamePeng
committed
ccc2af1
View commit details
Copy full SHA for ccc2af1
Browse repository at this point
prevent errors by setting 'image_min_tokens' to 256 if higher values are detected
alcoftTAO
committed
00e67b8
View commit details
Copy full SHA for 00e67b8
Browse repository at this point
Commits on Apr 5, 2026
updated code
alcoftTAO
committed
17366d3
View commit details
Copy full SHA for 17366d3
Browse repository at this point
Merge branch 'JamePeng:main' into lfm2.5-vl
alcoftTAO
authored
5fc4588
View commit details
Copy full SHA for 5fc4588
Browse repository at this point
feat(types): align with latest OpenAI OpenAPI spec (audio, structured outputs)
Show description for 1dc235f
JamePeng
committed
1dc235f
View commit details
Copy full SHA for 1dc235f
Browse repository at this point
Commits on Apr 4, 2026
Update Submodule vendor/llama.cpp b7ad48e..b863507
JamePeng
committed
a1a0c94
View commit details
Copy full SHA for a1a0c94
Browse repository at this point
docs: clarify enable_thinking compatibility for Gemma 4 models
Show description for 7bd7175
JamePeng
committed
7bd7175
View commit details
Copy full SHA for 7bd7175
Browse repository at this point
Update llama_types.py OpenAI OpenAPI Link
JamePeng
committed
6e99244
View commit details
Copy full SHA for 6e99244
Browse repository at this point
Update Submodule vendor/llama.cpp 277ff5f..b7ad48e
JamePeng
committed
73c3b06
View commit details
Copy full SHA for 73c3b06
Browse repository at this point
fix(Qwen35ChatHandler): Correct CHAT_FORMAT `tool_call` typo
Show description for 10a6abd
JamePeng
committed
10a6abd
View commit details
Copy full SHA for 10a6abd
Browse repository at this point
feat(chat_format): add Gemma 4 chat handler with multimodal and tool support
Show description for a5b4762
JamePeng
committed
a5b4762
View commit details
Copy full SHA for a5b4762
Browse repository at this point
Commits on Apr 3, 2026
Update llama_vocab_pre_type varriable
JamePeng
committed
7ef09e9
View commit details
Copy full SHA for 7ef09e9
Browse repository at this point
Update Submodule vendor/llama.cpp 12dbf1d..277ff5f
JamePeng
committed
2e27fd8
View commit details
Copy full SHA for 2e27fd8
Browse repository at this point
Merge pull request #103 from TAO71-AI/main
Show description for b751c83
JamePeng
authored
b751c83
View commit details
Copy full SHA for b751c83
Browse repository at this point
Implemented 'LFM25VLChatHandler'.
alcoftTAO
committed
aaf1922
View commit details
Copy full SHA for aaf1922
Browse repository at this point
fix Qwen3.5 chat template bugs
alcoftTAO
committed
9937acb
View commit details
Copy full SHA for 9937acb
Browse repository at this point
Commits on Apr 1, 2026
Sync llama : refactor llama_model_quantize_params to expose a pure C interface (#20346)
Show description for 24f2562
JamePeng
committed
24f2562
View commit details
Copy full SHA for 24f2562
Browse repository at this point
fix TypeError in _ggml.py
JamePeng
committed
7036ac3
View commit details
Copy full SHA for 7036ac3
Browse repository at this point
refactor(logger): migrate from llama_log_callback to ggml_log_callback
Show description for 225c7ad
JamePeng
committed
225c7ad
View commit details
Copy full SHA for 225c7ad
Browse repository at this point
feat(ggml): add support for ggml-base library and new function bindings
Show description for 08f15e7
JamePeng
committed
08f15e7
View commit details
Copy full SHA for 08f15e7
Browse repository at this point
Update Submodule vendor/llama.cpp 0fcb376..12dbf1d
JamePeng
committed
65d2750
View commit details
Copy full SHA for 65d2750
Browse repository at this point
Commits on Mar 31, 2026
Bump version to 0.3.34
Show description for a184583
JamePeng
committed
a184583
View commit details
Copy full SHA for a184583
Browse repository at this point
Update Submodule vendor/llama.cpp 2405d59..0fcb376
JamePeng
committed
a8cec00
View commit details
Copy full SHA for a8cec00
Browse repository at this point
Commits on Mar 29, 2026
docs(readme): add documentation for Assistant Prefill features
Show description for f9b5313
JamePeng
committed
f9b5313
View commit details
Copy full SHA for f9b5313
Browse repository at this point
feat(chat_format): added `assistant_prefill` to seamlessly continue responses
Show description for 3ead090
JamePeng
committed
3ead090
View commit details
Copy full SHA for 3ead090
Browse repository at this point
Update Submodule vendor/llama.cpp 59d8402..2405d59
JamePeng
committed
6d27d08
View commit details
Copy full SHA for 6d27d08
Browse repository at this point
Commits on Mar 28, 2026
Update README.md for Dynamic LoRA Routing & Control Vectors
JamePeng
committed
ee15771
View commit details
Copy full SHA for ee15771
Browse repository at this point
feat(LoRA): implement JIT dynamic LoRA routing and Control Vector injection
Show description for 57bfdd8
JamePeng
committed
57bfdd8
View commit details
Copy full SHA for 57bfdd8
Browse repository at this point
Commits on Mar 27, 2026
perf(context): debounce loras adapter clearing to prevent compute graph reallocation overhead
Show description for 82ef995
JamePeng
committed
82ef995
View commit details
Copy full SHA for 82ef995
Browse repository at this point
Previous
Next
You can’t perform that action at this time.