Title: Support 1 * 128 and 128 * 128 block-wise quant? · Issue #38 · NVIDIA/nvmath-python · GitHub
Open Graph Title: Support 1 * 128 and 128 * 128 block-wise quant? · Issue #38 · NVIDIA/nvmath-python
X Title: Support 1 * 128 and 128 * 128 block-wise quant? · Issue #38 · NVIDIA/nvmath-python
Description: In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quantize approach. I wonder do we have any pl...
Open Graph Description: In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quant...
X Description: In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quant...
Opengraph URL: https://github.com/NVIDIA/nvmath-python/issues/38
X: @github
Domain: github.com
{"@context":"https://schema.org","@type":"DiscussionForumPosting","headline":"Support 1 * 128 and 128 * 128 block-wise quant?","articleBody":"In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quantize approach. I wonder do we have any plan for support this approach?\n\nhttps://docs.nvidia.com/cuda/cublas/index.html#cublasltmatmulmatrixscale-t","author":{"url":"https://github.com/zfan2356","@type":"Person","name":"zfan2356"},"datePublished":"2025-08-04T08:21:37.000Z","interactionStatistic":{"@type":"InteractionCounter","interactionType":"https://schema.org/CommentAction","userInteractionCount":2},"url":"https://github.com/38/nvmath-python/issues/38"}
| route-pattern | /_view_fragments/issues/show/:user_id/:repository/:id/issue_layout(.:format) |
| route-controller | voltron_issues_fragments |
| route-action | issue_layout |
| fetch-nonce | v2:8a3a909d-cc21-404b-b74a-9d6eb51aad84 |
| current-catalog-service-hash | 81bb79d38c15960b92d99bca9288a9108c7a47b18f2423d0f6438c5b7bcd2114 |
| request-id | AC00:E0CBF:3B2B8:50273:6A61563F |
| html-safe-nonce | cd5f082917a024a499609013c728c89e7f28f93c45318411ecc38a79e519c1e3 |
| visitor-payload | eyJyZWZlcnJlciI6IiIsInJlcXVlc3RfaWQiOiJBQzAwOkUwQ0JGOjNCMkI4OjUwMjczOjZBNjE1NjNGIiwidmlzaXRvcl9pZCI6IjExMTg1NjUwMTkzMTM4NTQwMTUiLCJyZWdpb25fZWRnZSI6ImlhZCIsInJlZ2lvbl9yZW5kZXIiOiJpYWQifQ== |
| visitor-hmac | 43399cbafa6cd3d1bade867c72ab447d40ea99413a6df9719f2e4e9acef29fbc |
| hovercard-subject-tag | issue:3288501825 |
| github-keyboard-shortcuts | repository,issues,copilot |
| google-site-verification | Apib7-x98H0j5cPqHWwSMm6dNU4GmODRoqxLiDzdx9I |
| octolytics-url | https://collector.github.com/github/collect |
| analytics-location | / |
| fb:app_id | 1401488693436528 |
| apple-itunes-app | app-id=1477376905, app-argument=https://github.com/_view_fragments/issues/show/NVIDIA/nvmath-python/38/issue_layout |
| twitter:image | https://opengraph.githubassets.com/b31e41e2f6fa18e898fd460e803531ceb21406672fe430fd2bb5b41c696b9170/NVIDIA/nvmath-python/issues/38 |
| twitter:card | summary_large_image |
| og:image | https://opengraph.githubassets.com/b31e41e2f6fa18e898fd460e803531ceb21406672fe430fd2bb5b41c696b9170/NVIDIA/nvmath-python/issues/38 |
| og:image:alt | In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quant... |
| og:image:width | 1200 |
| og:image:height | 600 |
| og:site_name | GitHub |
| og:type | object |
| og:author:username | zfan2356 |
| hostname | github.com |
| expected-hostname | github.com |
| None | 35b919fdb8e6752d2ed95a144e761ca5e924557563e2ed0a5ebf6c4430cdeeb6 |
| turbo-cache-control | no-preview |
| go-import | github.com/NVIDIA/nvmath-python git https://github.com/NVIDIA/nvmath-python.git |
| octolytics-dimension-user_id | 1728152 |
| octolytics-dimension-user_login | NVIDIA |
| octolytics-dimension-repository_id | 788208729 |
| octolytics-dimension-repository_nwo | NVIDIA/nvmath-python |
| octolytics-dimension-repository_public | true |
| octolytics-dimension-repository_is_fork | false |
| octolytics-dimension-repository_network_root_id | 788208729 |
| octolytics-dimension-repository_network_root_nwo | NVIDIA/nvmath-python |
| turbo-body-classes | logged-out env-production page-responsive |
| disable-turbo | false |
| browser-stats-url | https://api.github.com/_private/browser/stats |
| browser-errors-url | https://api.github.com/_private/browser/errors |
| release | 1e22b08b60ccb67c3791f565bed248565260e611 |
| ui-target | full |
| theme-color | #1e2327 |
| color-scheme | light dark |
Links:
Viewport: width=device-width