Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Uh oh!
There was an error while loading.
Please reload this page
.
cosmasense
/
llama-cpp-python
Public
forked from
JamePeng/llama-cpp-python
Notifications
You must be signed in to change notification settings
Fork
0
Star
0
Code
Pull requests
0
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Pull requests
Actions
Projects
Security and quality
Insights
Commits
Breadcrumbs
History for
llama-cpp-python
llama_cpp
on
main
User selector
All users
All time
Commit history
Commits on Mar 13, 2026
Update mtmd_cpp API 20260313
JamePeng
committed
afd06ea
View commit details
Copy full SHA for afd06ea
View code at this point
Browse repository at this point
fix(Llama.generate): add explicit fallback context reset and expand `Llama.generate` docstrings
Show description for 6f81466
JamePeng
committed
6f81466
View commit details
Copy full SHA for 6f81466
View code at this point
Browse repository at this point
Commits on Mar 12, 2026
fix(sampling): prevent memory drift and hallucinations in logits view
Show description for aa38850
JamePeng
committed
aa38850
View commit details
Copy full SHA for aa38850
View code at this point
Browse repository at this point
Commits on Mar 11, 2026
feat(_ggml): extend ctypes bindings with more ggml constants, enums, and structs
JamePeng
committed
2bb5cd9
View commit details
Copy full SHA for 2bb5cd9
View code at this point
Browse repository at this point
Update llama.cpp API 20260312
JamePeng
committed
c420a2a
View commit details
Copy full SHA for c420a2a
View code at this point
Browse repository at this point
fix(chat_handler): fix tools and function calling in MTMDChatHandler.
JamePeng
committed
8641667
View commit details
Copy full SHA for 8641667
View code at this point
Browse repository at this point
Commits on Mar 10, 2026
perf(cache): optimize LlamaDiskCache I/O and fix LRU behavior
Show description for a36e8d2
JamePeng
committed
a36e8d2
View commit details
Copy full SHA for a36e8d2
View code at this point
Browse repository at this point
Sync examples : fix empty items in json_schema_to_grammar.py (#19968)
JamePeng
committed
8e8f36f
View commit details
Copy full SHA for 8e8f36f
View code at this point
Browse repository at this point
perf(cache): upgrade `LlamaRAMCache` to O(1) eviction and set `LlamaTrieCache` as default
Show description for 5c8e056
JamePeng
committed
5c8e056
View commit details
Copy full SHA for 5c8e056
View code at this point
Browse repository at this point
Commits on Mar 9, 2026
fix(core): disable swa_full for non-SWA models (sync llama.cpp upstream #20291)
Show description for 9acc507
JamePeng
committed
9acc507
View commit details
Copy full SHA for 9acc507
View code at this point
Browse repository at this point
Update llama.cpp API 20260310
JamePeng
committed
955ac33
View commit details
Copy full SHA for 955ac33
View code at this point
Browse repository at this point
fix(chat_format): fix namespace and variable shadowing of llama modules
Show description for 7efcb2b
JamePeng
committed
7efcb2b
View commit details
Copy full SHA for 7efcb2b
View code at this point
Browse repository at this point
fix(cache): fix namespace shadowing to prevent AttributeError
Show description for 5e285fe
JamePeng
committed
5e285fe
View commit details
Copy full SHA for 5e285fe
View code at this point
Browse repository at this point
Commits on Mar 8, 2026
feat(MTMDChatHandler): support audio inputs and fix interleaved media ordering
Show description for 964160d
JamePeng
committed
964160d
View commit details
Copy full SHA for 964160d
View code at this point
Browse repository at this point
Bump version to 0.3.32
JamePeng
committed
e7e1d48
View commit details
Copy full SHA for e7e1d48
View code at this point
Browse repository at this point
fix(sampling): pass seed to sampling context and remove global mutation
Show description for ad10cfd
JamePeng
committed
ad10cfd
View commit details
Copy full SHA for ad10cfd
View code at this point
Browse repository at this point
Commits on Mar 7, 2026
perf(hybrid): optimize multimodal single-turn and fix KV clear bug
Show description for fb3072d
JamePeng
committed
fb3072d
View commit details
Copy full SHA for fb3072d
View code at this point
Browse repository at this point
perf(hybrid): prevent expensive array slicing when cache is disabled
Show description for 850ed2e
JamePeng
committed
850ed2e
View commit details
Copy full SHA for 850ed2e
View code at this point
Browse repository at this point
perf(hybrid): bypass N-1 evaluation split if max_checkpoints is 0
Show description for 191e334
JamePeng
committed
191e334
View commit details
Copy full SHA for 191e334
View code at this point
Browse repository at this point
perf(hybrid): eliminate PCIe I/O latency for single-turn workflows
Show description for 7289b54
JamePeng
committed
7289b54
View commit details
Copy full SHA for 7289b54
View code at this point
Browse repository at this point
Commits on Mar 6, 2026
Bump version to 0.3.31
Show description for 2c191cb
JamePeng
committed
2c191cb
View commit details
Copy full SHA for 2c191cb
View code at this point
Browse repository at this point
Commits on Mar 5, 2026
fix(mtmd): remove OS-level log suppression to expose critical C++ errors
Show description for 118a1a8
JamePeng
committed
118a1a8
View commit details
Copy full SHA for 118a1a8
View code at this point
Browse repository at this point
fix(hybrid): implement N-1 checkpointing to support 1-token rollbacks
Show description for c03ce22
JamePeng
committed
c03ce22
View commit details
Copy full SHA for c03ce22
View code at this point
Browse repository at this point
Commits on Mar 4, 2026
refactor(mtmd): introduce omni-modal media pipeline with experimental audio support
Show description for b342f70
JamePeng
committed
b342f70
View commit details
Copy full SHA for b342f70
View code at this point
Browse repository at this point
Commits on Mar 2, 2026
Bump version to milestone version 0.3.30.
JamePeng
committed
41959f5
View commit details
Copy full SHA for 41959f5
View code at this point
Browse repository at this point
Commits on Mar 1, 2026
update chat handler
alcoftTAO
committed
2258973
View commit details
Copy full SHA for 2258973
View code at this point
Browse repository at this point
Merge branch 'JamePeng:main' into main
alcoftTAO
authored
db6ee16
View commit details
Copy full SHA for db6ee16
View code at this point
Browse repository at this point
refactor(mtmd): redesign multimodal pipeline for concurrent I/O and hybrid state management
Show description for 5bf6b6a
JamePeng
committed
5bf6b6a
View commit details
Copy full SHA for 5bf6b6a
View code at this point
Browse repository at this point
fix: Correct the mtmd vision check condition bug
JamePeng
committed
22df747
View commit details
Copy full SHA for 22df747
View code at this point
Browse repository at this point
refactor(chat_handler): extract MTMDChatHandler base class and Simplify the complexity of subsequent multimodal adaptation
Show description for 8b29b88
JamePeng
committed
8b29b88
View commit details
Copy full SHA for 8b29b88
View code at this point
Browse repository at this point
perf(eval): implement adaptive checkpoint intervals for hybrid models
Show description for 00da436
JamePeng
committed
00da436
View commit details
Copy full SHA for 00da436
View code at this point
Browse repository at this point
fix(eval): make context shift mathematically robust and architecture-safe
Show description for c27513d
JamePeng
committed
c27513d
View commit details
Copy full SHA for c27513d
View code at this point
Browse repository at this point
Add the memory_can_shift API to class LlamaContext
JamePeng
committed
667f433
View commit details
Copy full SHA for 667f433
View code at this point
Browse repository at this point
Commits on Feb 28, 2026
Merge branch 'JamePeng:main' into main
alcoftTAO
authored
3698316
View commit details
Copy full SHA for 3698316
View code at this point
Browse repository at this point
feat(eval): enable native context shift for hybrid/recurrent models
Show description for 4bfe5ed
JamePeng
committed
4bfe5ed
View commit details
Copy full SHA for 4bfe5ed
View code at this point
Browse repository at this point
Previous
Next
You can’t perform that action at this time.