gguf-quantization
GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.
[GGUFQuantizationllama.cppCPU
PoorRican
0
weights-and-biases
W&B: log ML experiments, sweeps, model registry, dashboards.
[MLOpsWeightsAndBiases
PoorRican
0
evaluating-llms-harness
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).
[EvaluationLMEvaluationHarness
PoorRican
0