Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
PrismML-Eng
/
llama.cpp
Public
forked from
ggml-org/llama.cpp
Notifications
You must be signed in to change notification settings
Fork
164
Star
865
Code
Issues
41
Pull requests
52
Discussions
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Security and quality
Insights
Actions: PrismML-Eng/llama.cpp
Actions
All workflows
Workflows
Release (Prism)
Release (Prism)
Build relocatable cmake package
Build relocatable cmake package
Check Pre-Tokenizer Hashes
Check Pre-Tokenizer Hashes
Check vendor
Check vendor
Code Style Checker
Code Style Checker
CodeQL
CodeQL
Convert PR to draft
Convert PR to draft
Copilot
Copilot
Copilot cloud agent
Copilot cloud agent
Copilot code review
Copilot code review
Show more workflows...
Management
Caches
All workflows
All workflows
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Showing runs from all workflows
will be ignored since log searching is not yet available
2,500+ workflow runs
2,500+ workflow runs
Workflow
Filter by Workflow
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching workflows.
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
cuda: add PTQ1_0 support
Update Operations Documentation
#39:
Commit
a294e8d
pushed by
bri-prism
34s
ptq1_0-cuda
ptq1_0-cuda
34s
View workflow file
cuda: add PTQ1_0 support
Python check requirements.txt
#37:
Commit
a294e8d
pushed by
bri-prism
2m 52s
ptq1_0-cuda
ptq1_0-cuda
2m 52s
View workflow file
cuda: add PTQ1_0 support
Python Type-Check
#218:
Commit
a294e8d
pushed by
bri-prism
3m 16s
ptq1_0-cuda
ptq1_0-cuda
3m 16s
View workflow file
cuda: add PTQ1_0 support
Check Pre-Tokenizer Hashes
#63:
Commit
a294e8d
pushed by
bri-prism
1m 4s
ptq1_0-cuda
ptq1_0-cuda
1m 4s
View workflow file
ggml: add PTQ1_0, ternary at group 128
Update Operations Documentation
#38:
Commit
7ce416f
pushed by
bri-prism
30s
ptq1_0-base
ptq1_0-base
30s
View workflow file
ggml: add PTQ1_0, ternary at group 128
Check Pre-Tokenizer Hashes
#62:
Commit
7ce416f
pushed by
bri-prism
1m 2s
ptq1_0-base
ptq1_0-base
1m 2s
View workflow file
ggml: add PTQ1_0, ternary at group 128
Python check requirements.txt
#36:
Commit
7ce416f
pushed by
bri-prism
2m 52s
ptq1_0-base
ptq1_0-base
2m 52s
View workflow file
ggml: add PTQ1_0, ternary at group 128
Python Type-Check
#217:
Commit
7ce416f
pushed by
bri-prism
3m 26s
ptq1_0-base
ptq1_0-base
3m 26s
View workflow file
cuda: prefetch the next PTQ1_0 mat-vec's weights into L2 from the last CTAs
Server
#429:
Pull request
#314
opened by
sb32445
Action required
sb32445:pr/ptq1-l2-prefetch
sb32445:pr/ptq1-l2-prefetch
Action required
View #314
View workflow file
cuda: prefetch the next PTQ1_0 mat-vec's weights into L2 from the last CTAs
Pull Request Labeler
#459:
Pull request
#314
opened by
sb32445
22s
22s
View #314
View workflow file
common : take the top-k straight from the logits when top-k starts the chain
Server
#428:
Pull request
#313
opened by
sb32445
Action required
sb32445:pr/sampler-topk-from-logits
sb32445:pr/sampler-topk-from-logits
Action required
View #313
View workflow file
common : take the top-k straight from the logits when top-k starts the chain
Pull Request Labeler
#458:
Pull request
#313
opened by
sb32445
17s
17s
View #313
View workflow file
server : reuse the buffers of evicted prompt checkpoints
Server
#427:
Pull request
#312
opened by
sb32445
Action required
sb32445:pr/server-ckpt-buffer-reuse
sb32445:pr/server-ckpt-buffer-reuse
Action required
View #312
View workflow file
server : reuse the buffers of evicted prompt checkpoints
Pull Request Labeler
#457:
Pull request
#312
opened by
sb32445
21s
21s
View #312
View workflow file
kv-cache : limit seq_rm to the used cell range
Server
#426:
Pull request
#311
opened by
sb32445
Action required
sb32445:pr/kv-seq-rm-bound
sb32445:pr/kv-seq-rm-bound
Action required
View #311
View workflow file
kv-cache : limit seq_rm to the used cell range
Pull Request Labeler
#456:
Pull request
#311
opened by
sb32445
16s
16s
View #311
View workflow file
cuda: gate (SwiGLU) fused PTQ1_0 mat-vec for 2-4 columns
Server
#425:
Pull request
#310
opened by
sb32445
Action required
sb32445:pr/ptq1-gate-up-fuse-mc
sb32445:pr/ptq1-gate-up-fuse-mc
Action required
View #310
View workflow file
cuda: gate (SwiGLU) fused PTQ1_0 mat-vec for 2-4 columns
Pull Request Labeler
#455:
Pull request
#310
opened by
sb32445
16s
16s
View #310
View workflow file
cuda: use the transposing concat kernel on all GPUs, not only on GB10
Server
#424:
Pull request
#309
opened by
sb32445
Action required
sb32445:pr/concat-transpose-all-gpus
sb32445:pr/concat-transpose-all-gpus
Action required
View #309
View workflow file
cuda: use the transposing concat kernel on all GPUs, not only on GB10
Pull Request Labeler
#454:
Pull request
#309
opened by
sb32445
16s
16s
View #309
View workflow file
cuda: smaller KV tile for the 8-column MMA flash attention config (head size 256)
Pull Request Labeler
#453:
Pull request
#308
opened by
sb32445
18s
18s
View #308
View workflow file
cuda: use the MMA flash attention kernel for GQA above 4 with quantized K/V on Ada
Server
#423:
Pull request
#307
opened by
sb32445
Action required
sb32445:pr/fattn-gqa-mma
sb32445:pr/fattn-gqa-mma
Action required
View #307
View workflow file
cuda: use the MMA flash attention kernel for GQA above 4 with quantized K/V on Ada
Pull Request Labeler
#452:
Pull request
#307
opened by
sb32445
15s
15s
View #307
View workflow file
cuda: PQ2_0 mat-vec kernel for 3-8 columns on Ada
Server
#422:
Pull request
#306
opened by
sb32445
Action required
sb32445:pr/pq2_0-multicol
sb32445:pr/pq2_0-multicol
Action required
View #306
View workflow file
cuda: PQ2_0 mat-vec kernel for 3-8 columns on Ada
Pull Request Labeler
#451:
Pull request
#306
opened by
sb32445
17s
17s
View #306
View workflow file
Previous
1
2
3
4
5
…
99
100
101
Next
You can’t perform that action at this time.