Need no worries, once I get enough money I will start funding small language model devs so we can get rid of the hands of big companies once and for all
appvoid
AI & ML interests
Recent Activity
Organizations
Covers:
- Hybrid architecture based on Qwen3.5
- Pre-training with 15B tokens
- Cost benchmark between H200 and B200
- Post-training with SFT + LoRA
- Full code and data, open source
With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline.
Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch
Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute.
You can now train & run LLMs on your AMD hardware
โข We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
โข Works on Windows, WSL, Linux
โข Train Qwen, Gemma on just 3GB VRAM
GitHub: https://github.com/unslothai/unsloth
Blog + Guide: https://unsloth.ai/docs/basics/amd
Yes, I got the script for ArithMark-3 and I can re-run it with no issues.
Mostly across harness details rather than a completed five-seed sweep. We have checked the checkpoint under repeated runs and small evaluation-pipeline variations, but I do not want to represent that as a formal variance estimate.
Your distinction is valid though, prompt formatting and answer parsing measure harness sensitivity, while multiple seeds under one frozen configuration might establish the actual noise band. I don't know if it will be published at all. Since is a little bit chaotic keeping base skills while improving them so our iteration method is barely reproducible.
Thatโs a fair point. At this scale, instruction tuning is definitely a capacity tradeoff rather than a free capability layer, so preserving the base modelโs strengths will be one of the main acceptance criteria.
Regarding the post-merge result, we have done additional validation runs and the performance appears directionally consistent, although I would prefer to publish the full repeated-evaluation results once the methodology and comparison conditions are finalized. At 90M, even small evaluation details can meaningfully affect the reported ranking.
The instruct version will be treated as a separate checkpoint rather than a replacement for the base model.
Can't wait to see what the community ๐ชdo with this! ๐๐๐
palmer-006 (90M)After 3 years of experiments, we are finally releasing our flagship tiny model: **palmer-006**.
If you are building for edge hardware, SBCs (Raspberry Pi, etc.), or low-power devices, this is for you. Inspired by Andrej Karpathy's idea of a self-contained "cognitive core," we wanted to see how much power we could pack into a sub-100M parameter footprint.
๐ง **How we "Palmerized" it:**
We believe in starting our experiments with the absolute strongest baseline possible.
1. Light fine-tuning on highly curated data
2. Model merging
3. Another light fine-tuning round
4. Adjusted Mamba for maximum token speed โก๏ธ
โ ๏ธ *Note: This is a foundational language model. It has not been instruction-tuned yet!*
Also, since this needs instruction tuning next to become a chat assistantโ**what dataset would you recommend we use for the instruct tune?**
---
๐ **Quick Links & Info:**
* **License:** Open for research, education, hobby, and modification! (For commercial use/hosted APIs, shoot an email to nosoyhackercodigo@gmail.com. *PS: Donators can claim a free commercial license!*)
* **Attribution:** Built using AI tech from the Technology Innovation Institute (TII).
Can't wait to see what you build at the edge. Let me know your prompt completions below! ๐
appvoid/palmer-006
We would like to clarify that SupraLabs has no affiliation, partnership, or connection whatsoever with "SupraLarps" or its members.
Please avoid interacting with their organization, repositories, or Spaces under the assumption that they are associated with us.
We are currently aware of the situation and have already contacted the appropriate channels to address it.
Thank you to everyone who continues to support SupraLabs. โค๏ธ
We are releasing Limen0.2B, a 222.5M-parameter base language model developed as a research platform for efficient pretraining and superword tokenization at smaller scales.
Limen0.2B was trained from scratch on 50B tokens and uses a compact 16K BoundlessBPE vocabulary. The project explores whether SuperBPE-style tokenization can remain effective in a substantially smaller model and vocabulary regime than those examined in earlier large-scale experiments.
The model also combines a deep-and-narrow transformer design with Exclusive Self-Attention, grouped-query attention, and tied embeddings. Its compact vocabulary reduces the embedding footprint and leaves a larger share of the parameter budget available to the transformer layers.
Despite its relatively modest training budget, Limen0.2B achieves competitive results for its scale across the reported language understanding, commonsense reasoning, and grammatical evaluation tasks. Comparisons with other compact models are provided as context rather than strict rankings, as their training data, token budgets, architectures, and evaluation settings differ.
The release includes the model weights, implementation, training configuration, checkpoint progression, and evaluation results, all under Apache 2.0.
UniversalComputingResearch/Limen0.2B
Technical feedback, independent evaluations, and further experiments with the model and tokenizer are welcome.
Cool stuff right there! Keep it up
It measures model performance on a variety of different tasks:
Language Completion
Common sense too
World Knowledge
Context Tracking
Quantitative
Logical Reasoning
Code Completion
Each has a different score and 1 overall score.
Submit your own model:
BananaMind/BananaMindBench-Leaderboard
Check it out:
BananaMind/BananaMindBench-Leaderboard
CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.
See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2
Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
because its my own and im currently training it
you should do more of that magic you did with hellaswag on BananaMind-2-Medium
BananaMind 2 pro... ok i know
Let's Go!!!
One is Kimi-K3, which I have heard of briefly. What's the other one? I hope it's a video generation model
You guessed right! One of them is the best biggest model. The other one is the best smallest.