logo

How Demo Day Care Visualized Benchmark Results with AI:GO

How Demo Day Care Visualized Benchmark Results with AI:GO

At JunctionX Korea 2026 hackathon, held at POSTECH in Pohang from August 21 to 23, Lablup and FuriosaAI posed a challenge: how far can you go with models you can hold in your own hands? Teams had to combine three models served on FuriosaAI RNGD servers (K-EXAONE-236B-NVFP4, gpt-oss-120b, and Qwen3-32B) using the Squad feature of Lablup's AI:GO, solve coding, math, and general benchmarks at the lowest possible cost, and build a visualization showing how their squad worked through the problems.

Demo Day Care team took first place in the track and the overall grand prize with an observability tool that visualizes benchmark scores, efficiency, and total token usage in 3D. We asked the team, made up of three undergraduates and one intern, how they used AI:GO's agent squad feature to design a squad that could solve benchmark problems in 48 hours.

* This interview was conducted on September 9.

1. About the Lablup–FuriosaAI Track

Q. Let's start with introductions. What role did each of you play at the hackathon?

Seokhyun Bae | I'm Seokhyun Bae, an undergraduate in my final semester. I put this team together, and my main job was to take on the challenge first, divide up the roles, and integrate everything.

Jaewon Lee | I'm Jaewon Lee, also an undergraduate. I handled all the visualization, and as the planner, I fed insights into how we built the frontend and backend. I did the UI/UX design as well.

Wonseok Yoo | I'm Wonseok Yoo, an undergraduate. I developed alongside the others, and whenever we had something working, I ran QA to find what needed fixing and either reported it or fixed it myself.

Rokyeon Kim | I'm Rokyeon Kim. I've graduated and am working as an intern now. I helped run the tests Seokhyun set up and mostly gave input on how to improve our prompts.

Q. Why did you choose the Lablup–FuriosaAI track out of the three?

Rokyeon Kim | With the other two tracks, any topic we picked would end up being some kind of public service, and we thought that would take domain knowledge and a genuine interest in that field. We had nothing there that would set us apart from other teams. On the other hand, Seokhyun, Wonseok, and Jaewon are all really into agents, and Jaewon is great at visualization. Track 3's topic matched our team's strengths perfectly, so we decided almost immediately.

Jaewon Lee | I wanted a track where we wouldn't stop at calling an API but could go deeper technically. Seokhyun is good with harnesses and knows ML, so we thought we could do well even if the track was hard. The fact that it was a problem you could solve with technology made it even more appealing.

Seokhyun Bae | We talked among ourselves about how, with topics that have clearly defined deliverables, every team would end up producing something similar. Honestly, we also didn't have a fresh, clearly useful idea to bring to the other tracks, and without that kind of purpose, we felt it would be neither meaningful nor fun. I'd been doing agent work and contributing to open source, and Jaewon had his strengths. We were a little worried that the deliverable wasn't clearly defined, but we were confident we could stand out with visualization as something concrete to show, and that we could score well too.

2. Building a Dashboard for People Who "Really" Use AI

ARGUS DASHBOARD

Q. Let's talk about your winning project. How did you come up with a 3D observability tool that shows score, efficiency, and tokens on three axes?

Jaewon Lee | Honestly, when I look at AI models, I think intelligence matters most. If you want good output, a bigger, heavier model is always better. But this competition's evaluation metrics included efficiency, total token usage, and score. As we kept running tests, sometimes the score was good and sometimes the efficiency was good, and the starting point was wanting the whole team to see that at a glance. So I built the graph with three axes. Even with nothing but points plotted in 3D space, you can tell right away that a run sitting in the positive direction on all three axes is efficient and performs well.

Jaewon Lee | People who use AI a lot keep checking benchmarks every time a new model comes out, thinking about which LLM they'd use if they were building a service. I wanted them to be able to easily compare results when they run similar tasks many times with slightly different variables, and to spot what's wrong the moment they look. In the end, all the teams converged on gpt-oss, right? I think if the other teams had used this tool, they'd have gotten there faster. We built it on the philosophy of making benchmark systems easy for people to make sense of. We started out solving our own problem, but it ended up becoming a more general-purpose visualization than that.

3. How Demo Day Care Designed Their Squad

Q. What did your final squad look like? We'd also like to hear how you got there.

Seokhyun Bae | Our final squad was one planner and one Universal Implementer. The planner was Qwen, and the implementer was gpt-oss. It took three iterations to get there. At first, we built a typical harness structure: a separate agent each for math, coding, and general tasks, with EXAONE on math, gpt-oss on coding and general, and Qwen as the planner, plus reviewer and repair agents. But problems came up at every stage, and above all, routing didn't work properly. So we removed the reviewer and repair agents and had gpt-oss check its own work once within the math, coding, and general tasks. Routing still wasn't solved, so in the end we went with a single universal implementer.

Q. How did you decide which model went where?

Seokhyun Bae | We tried different combinations. Qwen gave the best results as the planner, and when we flipped it, with Qwen as the implementer and gpt-oss as the planner, Qwen couldn't solve the problems. But running everything on gpt-oss wasn't the right answer once you factor in tokens, so we decided to use at least two models. If you look at series like "oh-my-agent," they use implementers with strong personas, and that always bothered me. Tasks aren't neatly divided into coding, math, and writing; they're far more multifaceted. So when I build subagents, I tend to strip out all the personas and use a single universal implementer. I've always preached a methodology I call "fan-out," and that's what gave me the confidence to make such a bold change.

Q. Could you tell us more about fan-out?

Seokhyun Bae | It's a name I came up with. My motivation for building my own harness and subagent allocation structure was actually cost. I once won a Cursor competition and got a lot of credits, and to make good use of them, I tried setting one model as the planner and another as the implementer, writing subagents into the Codex harness in code. At first I kept creating specialized agents. I'm a student, so I made a homework agent, a writing agent, a coding agent. That turned out to be way too constraining, and I thought it would be better to cut the cost of each agent reading its own system prompt.

The advantage of fan-out is that the agents make their own decisions. As the name suggests, it spreads outward: if you create three subagents, each of them creates three more, and so on. But that doesn't work if you have specialized agents with personas. A coding agent spawning another coding agent is fine, I guess, but a math agent spawning another math agent makes for a strange structure. Codex recently added a feature that lets you fork a conversation and assign it to a subagent, so I've built fan-out V2 around that and I'm using it now.

Q. The brief also asked teams to design an agent that decides when to give up. Did you consider a structure that backs off when it can't solve something?

Seokhyun Bae | We were going to, but didn't. We decided that adding one more agent cost too much. We also tested various ways to pick up partial credit, but some things, like the SymPy problems in the math task, just couldn't be solved reliably no matter what we tried. So we decided that what can't be solved can't be solved. It's not giving up; a wrong answer is just a wrong answer. And we moved on.

Wonseok Yoo | Normally, when you're developing or using AI, letting the agent say "I don't know" or "I give up" is really important for both researchers and developers, to save tokens and avoid wasted session time. But in this track, we felt that deciding when to give up and calling another agent to try again didn't add much. So I think it was good that we kept the structure simple, like Seokhyun's fan-out: split the task sensibly and let a highly capable general-purpose model handle as much as it can within its abilities.

Q. The benchmarks covered three areas: coding, math, and general. Did you use different prompts for each?

Seokhyun Bae | We didn't get rid of personas entirely. We put everything into a single subagent prompt: "If it's math, do this; if it's coding, do this; if it's a general task, do this." Basically one very long instruction. In some cases, the Qwen planner solved general tasks on its own and returned a fallback. There were plenty of constraints, like limits on context and input length, but we tried to keep things as straightforward as possible.

4. Impressions of the Three Models

The three models Lablup and FuriosaAI provided for the track varied in size, and each was priced differently. Because final scoring factored in token usage alongside benchmark scores, how sparingly a team used the expensive model became part of the strategy each team could play.

Q. What was behind the decision to drop EXAONE from your final squad?

Seokhyun Bae | The biggest reason was cost. We figured EXAONE was given to us for a reason, so we kept trying it in different ways. It was good at math, and in our first setup, the math agent was EXAONE. But under this track's pricing, we couldn't find a reason to justify that cost in our setup, so we eventually dropped it.

Wonseok Yoo | It was priced high, but the score didn't go up enough to match.

Q. What about Qwen and gpt-oss?

Seokhyun Bae | Qwen was a puzzling model. It was the best of the three at planning. We tried gpt-oss as the planner too, but results were best with Qwen, and it even handled some general tasks on its own at the planning stage. But when we had it do implementation, it struggled to solve the problems. It's the smallest of the three, so I think it ran into the limits of its size.

gpt-oss was a true all-rounder, just as you'd expect from GPT. It was consistently good at general tasks, math, and coding, so we kept it as the implementer. That said, it didn't offer enough of an edge as a planner, and filling every role with gpt-oss would have been too costly in tokens, so we left planning to Qwen.

5. AI:GO: Something New in Building Subagent Allocation Structures

Q. This was your first time using AI:GO. What did you like about it while building your squad?

Seokhyun Bae | Other harnesses don't have a feature for building the subagent allocation structure itself. Unless you're a developer, there's simply no way to assign subagents or put together a squad, so just being able to do that through a GUI is significant in itself.

Q. If one strength is that you don't need to be a developer to build a squad, how was it for someone less familiar with agentic tools?

Rokyeon Kim | Compared to the rest of the team, I know the least about harnesses and agentic skills. From that perspective, it was a bit hard to understand what the application actually does. There were quite a few tabs, and with so many supported features, I wasn't quite sure where to start.

Lablup's note) AI:GO began as part of a project to experiment with Lablup's AI-driven development methodology. Through JunctionX, Lablup started work on improving AI:GO's usability. Together with UI/UX experts and through more internal testing, we're working to build an environment where more people can leverage the performance of their own hardware and run agents more easily. Stay tuned for AI:GO 2.0, which we're hard at work on.

Q. If agents end up using apps more than people do, what should the screen look like?

Seokhyun Bae | Everyone's talking about that these days. That the only GUI left will be a chat window or a microphone.

Jaewon Lee | Whether it's a CLI or an app like Codex or Claude, when tasks keep running, people don't keep stepping in on what tools or agents can do on their own. Turning tools on and off is handled separately. So I think the center of the screen should be communication with the agent; in other words, what you're asking the agent to do.

6. The Thrill of Hackathons: Solving Problems You Didn't See Coming

Q. We should probably start with the hardest part of the 48 hours. Was there a moment of real crisis?

Seokhyun Bae | The first day. We spent it dividing roles and brainstorming, so we never got a chance to actually try AI:GO. When we got back to our motel room, I said, "I'll stay up all night on this, you guys get some sleep," and took it apart piece by piece on my own. Harnesses that verify answers in a Python kernel score really well these days, so I thought that approach would get us first place. But it turned out to be a completely closed environment. No external tools, no way to get my own code recognized inside it, and no way to move things in or out. I figured that out alone at around 4 a.m. There was nothing I could do, so I went to sleep, and the next morning when the team asked, I told them, "We're kind of doomed."

After some sleep, things started to come together the next day. I spent the whole day on the harness. I told Jaewon only that the API logs would be structured a certain way, said "Visualize all day today," and barely talked to him after that. We didn't communicate; we each just focused on our own part. Three of us worked as one team on the squad, and the visualization was split off. It was nerve-racking at the time, but looking back, that was when things went best, with each of us doing what we could do.

Q. Everyone at Lablup who saw your final presentation apparently thought you'd win. How did you prepare?

Jaewon Lee | The demo went through three rounds of visualization. First, we completed it with the data we were given. Second, we pulled in other teams' data from the leaderboard to compare benchmarks, because we wanted to reflect that in our model orchestration. Finally, we gathered only the data we had used ourselves and presented with that. We found out we'd be going on stage only about 10 minutes beforehand. I was originally going to present alone, but the other three understood how the squad worked better than I did, so we decided to do it together and pulled it together in a rush.

Rokyeon Kim | Something funny happened. After we submitted, we got a little ahead of ourselves and said, "We might win, so let's practice in English." That was right at the very end and we didn't spend long on it, but I think it left us just slightly more prepared for an English presentation than the other teams.

7. Teams in the Agentic Era: One Brain

Q. Junction is a hackathon that builds teams from planning, design, and development roles. But these days it's easy for developers to do work outside development and for non-developers to write code, so the lines between roles have blurred a lot. How did your team divide the work?

Wonseok Yoo | We talked a lot about methodology and roles even on the KTX down to Pohang. Jaewon had shown a real talent for UI/UX in the projects we'd done together before, so we decided he'd take that on by himself.

Jaewon Lee | Talent aside, there was actually something the four of us had agreed on. We needed to build a big MVP fast with vibe coding, and if features are tangled together, you can't grow the frontend quickly, especially when agents are the ones building it. So we decided in advance to implement everything in a decoupled structure.

Seokhyun Bae | Our conclusion was "Let's move like one brain." We're one person, and these computers aren't separate machines splitting up the work but a single computer. It's just that everyone's tokens and subscription plans are different, so we split it into four threads. So instead of saying "I'll do frontend, you do backend" and each implementing only our own piece on our own computer, we pooled all our capabilities first and then built. UI/UX still seems to need individual skill, but when it came to building the logic in particular, the three of us moved almost as one. It felt like running experiments on three laptops as batch 1, batch 2, and batch 3. Hearing other teams' stories afterward, a lot of them had friction, with people saying "The frontend is taking forever" or "I have nothing to do," and we'd agreed not to let that happen.

Jaewon Lee | To sum up, I think making communication the top priority from the planning stage cut out a lot of other inefficiencies. Having everyone look in the same direction and agree was the methodology that let us move as one team, in both speed and quality, in the age of vibe coding.

Q. In an era when agents do almost all the design and coding, if there's one thing the humans on a team have to keep hold of, what would it be?

Seokhyun Bae | I think it's the tech stack. Honestly, I'd say the reason I win awards at hackathons comes down to differences in tech stack. When you design an architecture, which framework to use, how to deploy, which component library to pick, every one of those choices is the stack. Non-developers just tell the agent to do it and let it do whatever it does, but being able to choose the right stack for each situation without trial and error is the one thing that separates good vibe coding from bad right now. It might sound like a developer's take, but that's exactly why I think you need to build up a lot of knowledge.

Wonseok Yoo | I'd pick intuition. When you work on a project, you write documents like PRDs and TRDs and check whether they still match the direction you're building in. With intuition, you can correct course more easily and push your direction forward with more conviction.

Jaewon Lee | I think it's ownership. No matter how many subagents you run or how much multitasking you do, you still need a gut sense of how everything within your area of responsibility is working. I don't think the transformer architecture of 2026 has structurally acquired the ability to plan ahead that people build up through 20 or 30 years of repeated practice. It doesn't perceive time the way people do. So planning and setting direction still rely heavily on human ability.

8. What the Hackathon Left Behind

Q. Could each of you sum up in a sentence or two what you took away from this hackathon?

Seokhyun Bae | Looking at the other projects more than our own, I realized how important purpose is. Whether it's vibe coding or anything else, I think that drive is what matters, and every time I go to a hackathon, I see a lot of ideas made just for the sake of having an idea, with no real purpose behind them. That's also why we chose Track 3. Without purpose, you don't get good results. I'm now more convinced than ever: have a solid drive, or don't do it at all.

Rokyeon Kim | For me, it's the importance of a leader. Lots of good opinions can come up, but in the end, it's the leader who cuts them down and picks one. Having a leader who could make those cuts quickly and point us in a good direction was what set our team apart.

Wonseok Yoo | I was reminded of how important fundamentals are. What gave us confidence in our direction was, in the end, the basic knowledge we already had.

Jaewon Lee | The ability to explain your own ideas. For a team to bring its thinking together, each person's ideas have to be explained properly so they can pass from person to person, and from person to AI. Without that ability, you have no choice but to leave everything to the agent. Even at the company level, whatever one person did with an agent, if they can't communicate it to others, I think the idea ends right there.

Q. Lastly, is there anything you didn't get to say?

Jaewon Lee | Whether it's GPU cluster companies, semiconductor makers, or companies building AI services on top of them, I really hope Korean companies all do well. Even setting AI technology aside, I believe infrastructure itself is national strength. To me, that's the purpose.

Seokhyun Bae | Having Track 3 at Junction made it so much fun. It was like a playground of a topic for people like me. The CEO and a lot of people from Lablup came by and answered all our questions, and going back and forth on whether to use EXAONE and figuring it out together was a ton of fun for me personally. The CEO really came across as someone who builds things, not someone who just runs a company.


We put a lot of thought into preparing this track. Participants ranged from students to working professionals, so we spent a long time discussing what level of difficulty to aim for and what the topic should be. After almost two days of not being able to settle on a topic, it was Lablup's CTO, who happened to be passing by, who handed us the idea of visualization, and that five-minute suggestion became the track's second challenge.

Through the hackathon and this interview, we got a clear sense of how much technical knowledge the people who chose this track bring and where their interests lie. Watching them take apart models and harnesses themselves, overhaul their squad structure several times in 48 hours, and name and use their own methodologies, we saw both their geeky side and their passion, and we came away deeply inspired as well. Our thanks to Seokhyun Bae, Jaewon Lee, Wonseok Yoo, and Rokyeon Kim of Demo Day Care for taking the time to talk with us. For the story of how Lablup prepared and ran the track, see the JunctionX Korea 2026 Hackathon Retrospective.


Interviewees Seokhyun Bae, Jaewon Lee, Wonseok Yoo, Rokyeon Kim (Demo Day Care)

Interviewer, Editor, and Photographer Jinho Heo (Lablup)

We're here for you!

Complete the form and we'll be in touch soon

Contact Us
lablup

Headquarter & HPC Lab

KR Office: 8F, 577, Seolleung-ro, Gangnam-gu, Seoul, 06143, Republic of Korea US Office: 3003 N First st, Suite 221, San Jose, CA 95134

  • facebook
  • youtube
  • Linkedin
  • GitHub

© Lablup Inc. All rights reserved.

We value your privacy

We use cookies to analyze site traffic, understand how visitors use our website, and improve our services. Necessary cookies for basic site functions are always active. Learn more

By clicking "Accept All", you agree to the storage of analytics cookies on your device. Click "Reject All" to keep only necessary cookies, or "Customize" to choose for yourself. You can change your settings at any time.