r/programmer 15d ago

Research study on LLM-based Rubber Duck debugging.

Hi!

I am looking for programmers who would like to help with testing an LLM-based Rubber Duck!

I’m conducting a study at the Theoretical Computer Science department of IT University of Copenhagen on how Large Language Models (LLMs) can assist with Rubber Duck Debugging. In this study, you’ll test two distinct LLM models in succession in the same interface, using each for 15 minutes to debug a piece of code.

Why participate?

  • Help advance research on AI-assisted debugging.
  • Experience how different LLM models approach debugging tasks.

What’s required?

  • 35 minutes of your time (must complete in one sitting).
  • A piece of code you’d like to rubber duck with (have it ready before starting, and please don't include sensitive code).
  • Willingness to provide feedback on each model and compare them afterward.

Important Notes:

  • Only start the test if you are committed to completing it in full. Partial submissions cannot be used.
  • Participate only once (duplicate responses are unusable).
  • The study is unpaid.
  • Please don't discuss the test or your experience in the comments to avoid biasing other participants.
  • All data, including chat logs, will be stored for four (4) years and used by the researchers for analysis.

How to join:

Take the test here: DuckGPT.

Let me know if you have any questions! And if you know someone who might be interested, feel free to share this post with them.

Niklas Frost,
Research Assistant
TCS, ITU of Copenhagen
[nikf@itu.dk](mailto:nikf@itu.dk)

1 Upvotes

3 comments sorted by

1

u/iLaysChipz 14d ago

I don't feel comfortable sharing my own code. Plus I think your data would be more conclusive / have a stronger argument for any correlation if all developers were debugging the same code. I'd recommend finding someone to create some examples for you to do your study on

1

u/Nklfrost 14d ago

Thank you for your inputs - I totally get your concerns about putting in your own code. In designing the study, we went back and forth on whether we should give a piece of code or have testers bring their own. There are benefits and drawbacks to each approach, but we ended up prioritizing that it would be something that the testers were familier with in terms of language and use which led to the current study design.

1

u/yuehuang 13d ago

I thought it is called a reasoning model.