From Blink to FreeRTOS on the SG2000 Little Core
2026-10-05
The SG2000’s little C906 core is sold as a bare-metal Arduino target. You write setup() and loop(), and it blinks an LED.
That works. For one job, it is the right tool.
But as soon as you ask the core to do more than one thing — with any timing discipline — you rediscover every reason operating systems exist.
So the real question is not “what does FreeRTOS do?” It is:
Why do we want an OS at all — and why do we still want one on a tiny real-time core?
The SG2000’s second core is not a Linux CPU. nproc on the board reports 1.
| Property | Value |
|---|---|
| Core | T-Head C906L (RISC-V RV64), ~700 MHz |
| Cache | None (the big C920 has L1/L2) |
| Privilege | Starts in M-mode; no MMU page tables, no PMP |
| Memory | 2 MB carveout at 0x9fe0_0000 |
| Loaded by | remoteproc (/sys/class/remoteproc/remoteproc0) |
| Role | bare-metal / RTOS service core alongside Linux |
Because it runs in M-mode with no translation, a load or store is a physical access. Nothing is isolated — not from the SoC, not from Linux.
The vendor sophgo:SG200X Arduino core gives you exactly two functions:
setup() runs once; loop() runs forever — one thread of control.pinMode/digitalWrite already hide the pinmux and register writes.delay() is a busy-wait: the core spins and does nothing else.Build it with arduino-cli compile --fqbn sophgo:SG200X:duos; it links at 0x9fe0_0000 and is loaded via remoteproc.
The superloop is a complete, correct program. It is also a mental model with hard limits:
Everything we do for the rest of this deck is an answer to one of these:
| # | The question | The OS service that answers it |
|---|---|---|
| 1 | Do two things at once | tasks + scheduler |
| 2 | Meet a deadline without spinning | timers, blocking delays, tick |
| 3 | Run the important thing first | preemptive priority scheduling |
| 4 | Share data safely | critical sections, mutexes, queues |
| 5 | Write portable, reusable code | drivers, HAL, standard APIs |
| 6 | Cooperate with Linux | message queues / IPC |
An operating system is not a product. It is a bundle of answers to these questions, packaged so every program doesn’t reinvent them.
Goal: blink pin 7 at 1 Hz and pin 13 at 4 Hz — independently.
Naive superloop:
The two activities serialize: pin 13 only runs after pin 7 finishes, so neither rate is right. Blocking delay() is the enemy.
The OS answer: let each activity be a task with its own stack and its own notion of time; a scheduler interleaves them.
A task = an independent flow of control with:
static void blink_task(void *arg) {
const Blink *b = (const Blink *)arg;
for (;;) {
digitalWrite(b->pin, HIGH); vTaskDelay(b->half);
digitalWrite(b->pin, LOW); vTaskDelay(b->half);
}
}
xTaskCreate(blink_task, "fast", 256, &fast, 2, NULL);
xTaskCreate(blink_task, "slow", 256, &slow, 1, NULL);
vTaskStartScheduler();The scheduler gives the CPU to whichever ready task should run. loop() is gone; the OS owns the program counter now.
Goal: toggle a pin every 1 ms — precisely — while doing other work.
Naive approach: poll a counter or call delay(1).
delay() wastes the CPU and drifts under load. Polling in loop() has jitter proportional to how long the rest of the loop takes.
The OS answer: a hardware timer raises an interrupt at a fixed period; vTaskDelay / vTaskDelayUntil let a task sleep until a deadline, yielding the CPU in the meantime.
The C906 raises a machine-timer interrupt when a counter reaches a compare value:
mtime — free-running 64-bit countermtimecmp — compare register; interrupt fires when mtime ≥ mtimecmpmcause = 0x8000_0000_0000_0007 (interrupt, machine timer)FreeRTOS turns this into a periodic tick:
Each tick it re-arms mtimecmp and decides whether to switch tasks.
A task can then say “wake me in 10 ticks” and block — the CPU runs other tasks until the deadline. No spinning.
Goal: a safety-critical control loop must preempt a logging task.
In a superloop, ordering is lexical — whoever you wrote first runs first. There is no way to say “this matters more.”
The OS answer: priorities + preemption. - Every task has a priority. - The scheduler always runs the highest-priority ready task. - When a higher-priority task becomes ready, it preempts the running one at the next tick (or immediately, for an ISR-driven unblock).
This is what “real-time” actually buys you: a bounded, predictable delay before the important work runs — not raw speed.
Goal: loop() and an ISR both touch a counter.
If the ISR fires between loop()’s load and store, one increment is lost. This is a race condition — and it is a hardware problem, not a language one.
The OS answer: primitives that make a region atomic or hand data over safely: - critical sections (mask interrupts briefly) - mutexes (mutual exclusion, with priority inheritance) - queues / semaphores (move data and signal, race-free)
Look again at the Arduino sketch:
None of those names the SoC. The Arduino core maps them to VIVO_D3, the pinmux at 0x03001150, and SWPORTA_DR at 0x03021000.
That is already an OS-style abstraction layer — a tiny one. The OS answer generalizes it: - device drivers behind a stable interface, - a HAL so the same program runs on a different board, - standard APIs (tasks, queues, timers) that outlive the hardware.
Portability is not a luxury. It is what lets a program outlive the chip it was written for.
The little core is not alone. It shares DRAM with the C920 running Linux.
Raw cooperation means: - shared memory in the carveout, - a mailbox block at 0x0190_0000 to raise inter-core interrupts, - virtio rings (vdev0vring0/1, vdev0buffer) for messages.
Doing this by hand — polling a magic address, hand-rolling a ring buffer, getting the fences right — is exactly the kind of code an OS should own.
The OS answer: an IPC layer. FreeRTOS provides queues and task notifications; the vendor port layers RPMsg and the rtos_cmdqu command queue on top, so Linux sends a message and a task wakes up to handle it.
| You want to… | Without an OS | With an OS |
|---|---|---|
| do two things | interleave by hand, one PC | tasks + scheduler |
| hit a deadline | spin or poll, drift | timer tick + blocking sleep |
| prioritize | reorder your code | priority + preemption |
| share data | hope the race doesn’t happen | critical sections, mutexes, queues |
| port to a new chip | rewrite every register write | drivers + HAL + standard API |
| talk to another core | invent your own protocol | IPC / message queues |
Every “OS feature” is a named, reusable answer to a coordination problem you would otherwise solve badly, once per project.
A common intuition: “real-time means bare metal, because an OS adds overhead.”
Real-time means deterministic. The requirement is a bounded time to respond, not the smallest average time.
An RTOS increases determinism: - a fixed tick gives a known scheduling granularity, - priority preemption bounds the delay before urgent work runs, - blocking primitives remove unbounded polling loops, - bounded interrupt latency and critical sections.
The danger to real-time is unbounded behavior — and hand-rolled superloops are full of it.
If OSes are good, why not put the real OS on the little core?
Because Linux cannot run there:
| Requirement | Little C906L |
|---|---|
| MMU / page tables | none — M-mode, physical addressing only |
| Privilege separation | no S/U-mode setup; no isolation |
| Memory | 2 MB carveout (Linux wants far more) |
| Cache | none |
| SMP | nproc = 1; it’s a service core, not a Linux CPU |
If you want OS services on this core, you need the smallest OS that fits: a real-time kernel — FreeRTOS.
On the Duo S, the little core is there to serve the system:
A service provider needs two things an OS provides:
An RTOS is what turns “a core that can do anything” into “a component you can rely on.”
An OS is not free. It costs:
For one trivial job — Blink, a single ADC loop, a fixed waveform — bare metal is simpler, smaller, and correct.
The engineering skill is not “always use an OS.” It is recognizing the moment the coordination problems outgrow the superloop.
The vendor SG2000 port creates its own tasks in main_cvirtos() before calling vTaskStartScheduler() — e.g. a high-priority CMDQU task that handles Linux mailbox commands.
Bare metal
CPU is busy the whole time. Nothing else can run.
Same shape, opposite meaning.
delay()consumes time;vTaskDelay()surrenders it. This one change is most of what an RTOS buys you.
Every tick, the machine-timer interrupt fires:
mtvec holds the trap handler address (set before vTaskStartScheduler).mie.MTIE enables the machine-timer interrupt.A context switch is a save + restore of a running task’s state:
Save (into the task’s TCB / stack): - all integer registers, - mepc (where to resume), - mstatus (interrupt-enable state).
Restore (from the next task): - its registers, mepc, mstatus, - then mret returns to the next task.
FreeRTOS pre-builds a synthetic frame in pxPortInitialiseStack() so the first task can be started by a normal context-restore. xPortStartFirstTask() simply restores it and mrets.
The OS creates the illusion of many CPUs by swapping register state on one real core.
| Primitive | Use it for |
|---|---|
| Queue | move data between tasks; also signals a waiter |
| Binary semaphore | signal an event (e.g. from an ISR) |
| Counting semaphore | track N resources / events |
| Mutex | mutual exclusion; supports priority inheritance |
| Task notification | fast, lightweight per-task signal |
| Stream/message buffer | byte/record stream between task and ISR |
ISRs cannot block, so they use the ...FromISR() variants (xQueueSendFromISR, vTaskNotifyGiveFromISR) to wake a task that does the slow work — deferred interrupt processing.
In a superloop there is one stack. With tasks:
heap_4 is the usual allocator in this port.This is the concrete price of concurrency, and the reason “just add tasks” is not free on a memory-starved core.
We solve the same problem — blink two pins at different rates — four ways, each fixing what the previous one broke:
| Rung | Mechanism | Fixes |
|---|---|---|
| 1 | delay() superloop |
— (baseline) |
| 2 | millis() state machine |
concurrency without blocking |
| 3 | timer ISR | precise timing, decoupled from loop() |
| 4 | FreeRTOS tasks | priority, blocking, clean structure |
Each rung is more capable — and more complex. Watch the trade.
delay() SuperloopSimple. Correct for one job. Breaks with two: the activities serialize and neither runs at its intended rate.
millis() Non-Blocking State Machine#define LED_A 7
#define LED_B 13
static uint32_t tA, tB;
void setup() {
pinMode(LED_A, OUTPUT);
pinMode(LED_B, OUTPUT);
}
void loop() {
uint32_t now = millis();
if (now - tA >= 500) { tA = now; digitalWrite(LED_A, !digitalRead(LED_A)); }
if (now - tB >= 125) { tB = now; digitalWrite(LED_B, !digitalRead(LED_B)); }
}Both rates run independently — no blocking. But this is cooperative: a long step anywhere still delays everyone, and there are no priorities.
#define LED_A 7
volatile bool tick_a = false;
extern "C" void trap_handler(void) {
// mcause == machine-timer? re-arm mtimecmp, then:
tick_a = true;
}
void setup() {
pinMode(LED_A, OUTPUT);
// csrw mtvec, trap_handler; program mtimecmp; set mie.MTIE
}
void loop() {
if (tick_a) { tick_a = false; digitalWrite(LED_A, !digitalRead(LED_A)); }
}Exact period, decoupled from loop(). But: the ISR must be short, and tick_a is a shared variable — now we need atomicity (Q4).
Illustrative — the vendor core’s timer API may differ; see the note.
#include <FreeRTOS.h>
#include <task.h>
typedef struct { uint8_t pin; TickType_t half; } Blink;
static void blink_task(void *arg) {
const Blink *b = (const Blink *)arg;
pinMode(b->pin, OUTPUT);
for (;;) {
digitalWrite(b->pin, HIGH); vTaskDelay(b->half);
digitalWrite(b->pin, LOW); vTaskDelay(b->half);
}
}
void main_cvirtos(void) {
static const Blink fast = { 7, pdMS_TO_TICKS(125) };
static const Blink slow = { 13, pdMS_TO_TICKS(500) };
xTaskCreate(blink_task, "fast", 256, (void *)&fast, 2, NULL);
xTaskCreate(blink_task, "slow", 256, (void *)&slow, 1, NULL);
vTaskStartScheduler();
}Two independent, blocking, prioritized flows. No state machines; each task reads like the single-job sketch. The OS supplies the concurrency.
Rung 1 delay |
Rung 2 millis |
Rung 3 ISR | Rung 4 FreeRTOS | |
|---|---|---|---|---|
| Two independent rates | ✗ | ✓ | ✓ | ✓ |
| Blocks the CPU | yes | no | no | no |
| Priority | ✗ | ✗ | ✗ | ✓ |
| Determinism | poor | load-dependent | good | good |
| Shared-data safety | n/a | n/a | manual | provided |
| RAM cost | minimal | minimal | small | per-task stack |
| Code complexity | lowest | medium | medium | higher upfront |
Capability rises, complexity rises, and so does the footprint. That is the whole trade-off, on one slide.
The Sophgo SDK ships a FreeRTOS little-core port under freertos/cvitek/:
cvirtos.bin,0x9fe0_0000,main_cvirtos() → create tasks → vTaskStartScheduler().The Arduino sketch and the RTOS image are the same delivery mechanism. FreeRTOS is simply a richer program loaded the same way.
Cooperation happens through real hardware:
The vendor rtos_cmdqu layer packages this as a command queue; the high-priority CMDQU task dispatches commands to the right FreeRTOS task.
Compare the two worlds for “make the little core do X”:
Raw, no OS
With an RTOS
The OS is what turns a clever bare-metal hack into a system.
The path you already have:
The examples BlinkMillis, BlinkTimerISR, and BlinkRtos in this repo walk the four-rung ladder. (Rungs 2–4 are reasoned, not yet board-tested.)
Lab ladder (climb it on the DuoS):
Blink and Blink4; measure the pad with the /dev/mem sampler.BlinkMillis; prove two rates coexist.millis().Discuss:
CS 4250 · Why an OS? — From Blink to FreeRTOS