

Milka chocolate the same
Same, but mostly because they went from 100g to 98g to 95g to 90g, those greedy fucks.


Milka chocolate the same
Same, but mostly because they went from 100g to 98g to 95g to 90g, those greedy fucks.
Curious about the quant tho.
Q8 from unsloth.
Something like Qwen3.5-122b
My go to model for knowledge. Definitely much faster at Q5 but it lacks the tool calling quality of the Qwen3.6 models. Really hoping we see a Qwen3.6-122b soon…
About 200 t/s prompt processing and 10-20 t/s with MTP.
Greatly depends on the task, predictable things like code generates at 18-20 t/s. Creative writing more like 10-17 t/s.
Yes, I got a Strix Halo machine before the RAM price hike and use it to run all my ML stuff on it.
Currently using llama-swap with llama.cpp/ComfyUI and opencode/Open WebUI as frontend.
I’m running Qwen3.6-27b, Voxtral Mini 4b, Piper and Qwen Image. Also, some embedding and reranking models.
I use them for:
If you have trouble with outgoing mails, you can use a hybrid approach.
Receive mails directly to your server but use a mail service to relay your outgoing mails. Configuration for that is very simple in mailcow and there are a few dozen (free) transactional email providers (e.g. Scaleway).
That way you can keep receiving your mails privately and only have to give up some privacy when sending mails.


According to their now deleted Reddit account, it will use a custom Proton build on Intel and Zen 4 CPUs. For every other CPU you will need a kernel module.


As a side note, Qwen3.6-27B is much more capable than Qwen3.6-35B, even though it is much slower.
https://huggingface.co/unsloth/Qwen3.6-27B-GGUF
For coding tasks where you don’t mind waiting, you should be able to barely squeeze in the 8-bit quantized version with 32 GB RAM + 8 GB VRAM and have a pretty competent local model. 4-bit quants work but they have issues with complex tool calls.
If you use the MTP branch of llama.cpp (and a suitable model) you can even double or triple your token generation speed: https://github.com/ggml-org/llama.cpp/pull/22673
For easier tasks, disable reasoning for instant responses.


That looks pretty good. Looks like Portainer is getting replaced this weekend.


I share photos and videos through my Matrix/Signal bridge all the time just fine.
Calls don’t work but that’s about it.
Do you actually train the LLM or use RAG? I have been looking for a local LLM + Wikipedia RAG solution for a while now.
For now I just have kiwix-serve + searxng doing a simple search but the Kiwix search is…questionable.


I wrote an application which runs on my server and monitors my favorites on Tidal/Deezer/Qobuz. It downloads them in bulk whenever I have a premium account with one of them. Usually I purchase a month of premium every few months, at which point I get nice clean FLACs for local use.
The FLACs are moved to Jellyfin and I stream them using Finamp, which also supports transcoding, so I keep 128 kbps Opus files for offline playback and stream the raw FLAC files when bandwidth is no concern.
I have amassed a huge music library over the last decades, so even if all streaming websites go under tomorrow, I have enough music locally to last me a lifetime.


What do you want to host?
If it’s a simple text-only/Javascript website, you can host through Codeberg/Gitlab/Github on a custom domain free of charge.


If we’re talking about online editing, Collabora has web editors based on LibreOffice but with a modern UI: https://www.collaboraonline.com/
They are really great and can be self hosted (e.g. with Nextcloud).
For offline editing, as already mentioned, LibreOffice has an optional ribbon UI and OnlyOffice looks pretty modern as well.


CoreELEC can do it on Dolby Vision certified devices if you’re looking for a open source solution.


Is it a Surface laptop?
Fedora 43 with the Rawhide kernel.
gpt-oss is pretty much unusable without custom system prompt.
Sycophancy turned to 11, bullet points everywhere and you get a summary for the summary of the summary.
Of course, self hosted with llama-swap and llama.cpp. :)
I have a Strix Halo machine with 128GB VRAM so I’m definitely going to give this a try with gpt-oss-120b this weekend.
Awesome, did anyone test the VoLTE on a Fairphone 5 already?
I have a spare one laying around that I would love to get fully working with postmarketOS.