TIME-OUT

Vancouver, BC · Product, strategy, technology

Hit a letter three times. Hold Shift to attract the loose ones.

I'm Abhi. I turn rough problems into useful products across AI, planning, and new ventures.

Your desk titleproduct + code + judgment

Enter the work ↓
0 / 10 released

Three projects, opened up

The work behind
the work.

Scroll slowly. Each chapter has something to operate.
BoardroomAgent infrastructure01 / 03

One desk for an entire agent team. 

Boardroom turns parallel coding agents into an operating system: assign work, isolate it, watch it run, resolve conflicts, and review the result without stitching together terminal windows by hand.

My role
Product, interaction, architecture, full-stack build
Core stack
Next.js, TypeScript, SQLite, node-pty, SSE, Git worktrees
What shipped
Self-hosted agent OS, review queue, plan engine, nine-tool MCP server
Brief on the tableShip passwordless auth0/5 tasks assigned
Discarddrop task

Drag at least three cards onto the agents. Double-click a card to duplicate it.

assigned0finished0tokens0spend$0.00tests0

The problem

More agents create more coordination work. Parallel branches collide, context disappears between sessions, and “done” still needs a human to find the diff, test it, and decide what ships.

The system

Named personas run Claude, Codex, Hermes, or OpenCode inside isolated Git worktrees. Plans can run sequentially or in parallel; completed work enters a push-request queue, and a resolver agent handles merge conflicts.

The hard part

The UI is only the surface. Underneath it are persistent sessions, PTY process control, live SSE logs, a four-second dispatcher, branch lifecycle management, and enough recovery state to survive an agent failing halfway.

Explore the repository
145passing tests
4agent runtimes
30+REST endpoints
9MCP tools
PocketLogConsumer finance02 / 03

Receipts in. Clarity out. 

PocketLog removes the tiny decisions that make expense tracking fall apart. Scan the thing already in your pocket; the product turns it into useful, shared financial context.

My role
Product direction, experience design, application build
Core loop
Scan receipt, verify extraction, update budget, share the picture
What shipped
Mobile-first expense tracker with receipt scanning and family budgets
POCKETLOG LABTurn paper into a useful decision.

Drag the receipt into the scanner.

YOUR POCKET
DROP RECEIPT HERECAMERA FRAME
inputpaper receiptmodelvision extractionresultloose

The behavior problem

Budget tools fail when logging costs more attention than the purchase. PocketLog starts from the receipt, then extracts merchant, amount, and category so the useful habit takes one action instead of six.

The product loop

Capture feeds a live budget, the budget reveals spending patterns, and family sharing turns private bookkeeping into a shared picture. The value arrives immediately, then compounds with every scan.

The design judgment

AI stays backstage. People see an editable result rather than a chatbot: a faster default, clear confidence, and an obvious place to correct the machine when it guesses wrong.

Open PocketLog
1 scanfrom paper to entry
3 fieldsfilled automatically
livebudget feedback
sharedfamily context
SQL-R1Model distillation03 / 03

The best model was hiding at iteration fifty. 

A weekend experiment became a lesson in evaluation: more data was not always better, DPO actively hurt, and the default final checkpoint concealed the strongest result.

My role
Experiment design, trace pipeline, training, evaluation, write-up
Core stack
Python, MLX-LM, LoRA, DeepSeek-V4-Pro, Qwen2.5-Coder-7B
Proof standard
200 held-out BIRD questions, deterministic SQL execution match
CHOOSE A TRAINING PATHDon’t trust the last checkpoint.
ITER 5055.0%
THE CHECKPOINT
ITER 5055.0%best measured checkpoint
THE MODEL IS MEMORIZING
ITER 0ITER 300
generalization peak55.0% pass@1

The useful path rises early, peaks at iteration 50, then gives accuracy back as training continues.

The question

Can a seven-billion-parameter local model absorb SQL reasoning from a much stronger teacher cheaply enough to make the experiment reproducible on Apple Silicon?

The pipeline

DeepSeek-V4-Pro generated traces for BIRD questions. Only execution-correct SQL survived. A rank-16 LoRA adapter trained on 708 accepted traces, then every checkpoint was scored by deterministic execution match.

What the failures taught

Combining a weaker teacher doubled the data but underperformed. DPO fell to 24–27.5%. The winning iteration-50 checkpoint reached 55%; the automatic last save had already overfit back to 49%.

Read the experiment log
+7.5 ppexecution accuracy
$5.83teacher-model spend
708accepted traces
14 hlocal GPU time

The operating system

Business school taught me to ask why. Building taught me to finish. I work best where product judgment, technical curiosity, and a bias toward shipping overlap.

I'm Abhi Poluri, an SFU BBA student in Vancouver. My projects range from local-first AI tools to planning products and mobile apps.

I also built a profitable retail-arbitrage operation from zero startup capital. Different medium, same instinct: understand what matters and close the loop.

Based in
Vancouver, Canada
Studying
Strategy & entrepreneurship
Looking for
Product and technology internships

Word forge

Build the invitation.

Load all three materials into the forge. They're the combination I bring to internships, collaborations, and hard product problems.

Drop materialDrop materialDrop material
0 / 3 LOADED

Feed the machine.

Use Product, Strategy, and Code

There are two keyboard secrets hidden on this page.