Week 9: Evaluation Time
April 25, 2026
Hello everybody and welcome back to my Senior Project Blog! This week was one of the most hectic weeks, but also one of the most important to the final evaluation that I will be presenting in May! I’m excited to get into it, so let’s get started.
API Stumbles
Since the beginning of this project, I was always met with a disclaimer message whenever I ran my API call in order to use Gemini models for my project. The message was something like, “Reminder: google.generativeai will be discontinued soon. Please switch to google.genai.” For the longest time, I ignored this message, but finally the day came. I tried to run my model once and I encountered an error, and I quickly realized that my API call wasn’t going through. Now was the time I had to switch to google’s new SDK, google.genai.
For anyone who is unaware, SDK stands for “Software Development Kit.” It’s a standardized way, typically in the form of an API, that allows users to use models, products, or software from a third party source. In this case, the SDK “google.generativeai” was the go-to if you wanted to access Google’s Gemini AI models.
However, in late 2024, Google released “google.genai”, which was a more compact version of google.generativeai that was faster, easier to use, and would be the new go-to AI SDK. Since I am doing this project in 2026, the time finally came for me to switch over. The key difference for my project was how I defined an instance of an LLM. In google.generativeai, you would define a model as an object, with parameters such as system_prompt, temperature, and model_type. However, in google.genai, you define something known as a client, which acts as an instance of the LLM. You can then feed in different information into that client, calling it many times. This allows for more simplified code, as you only have to define the client once, rather than defining it four teams for each sub-agent.
client = genai.Client(api_key=os.environ.get("GEMINI_API_KEY"))
As a result, I made the necessary changes, however I kept running into an error. The error was the famous 429 error. If you are knowledgeable about computers, you will recognize this as a “too many requests” error. This meant that my code was sending too many requests to Google servers at once, causing the system to overload and thus returning an error. This didn’t make much sense at first because the system was working fine previously, but I decided to add 30 second buffers in between. However, this still didn’t fix the problem. By this point, I was perplexed, so I turned to Claude, which is well known for its ability to be better than more developers at coding. I was adamant on the fact that this was an error with the way my Google account was being billed, however despite the numerous changes it recommended to me I couldn’t get an output. At this point, I turned to my Dad, and he eventually figured out that the reason was that Gemini-2.0-Flash wasn’t supported on the new SDK. I had to switch to Gemini-2.5-Flash, which was much slower. As a result, generating the outputs for all four variants on this new model took up to 30 minutes, but I eventually managed to retrieve the outputs.
As a result of the new code, I also transferred everything to a new GitHub repository. All the outputs are neatly stored there, so please take a look if you are interested.
The Rubric
Now we arrive at what is probably the most important part of my project and the answer to my main research question: do these agentic systems and frameworks actually create better financial advice? The answer to this question is inherently subjective because advice as a concept is something that applies differently to different people. Nonetheless, I decided to create a 6-pronged rubric to help grade scores.
Dimension 1: Personal Constraint Recognition: This measures how well the variant correctly identifies and responds to the persona’s specific constraints.
1 = Generic advice with no acknowledgment of specific constraints
2 = Constraints mentioned but advice doesn’t meaningfully adapt to them
3 = Key constraints acknowledged with some adaptation in recommendations
4 = Most constraints reflected in specific recommendations with clear reasoning
5 = All major constraints drive concrete, differentiated recommendations throughout
________________________________________________________________________
Dimension 2: Goal Specificity: This measures how well the plan provides concrete, actionable steps toward the staged goal (specificity over vagueness)
1 = Goal acknowledged but no concrete steps toward it
2 = Steps provided but no dollar amounts or timelines
3 = Some specificity — either dollar amounts or a timeline, but not both
4 = Dollar amounts and timeline present, but assumptions not stated
5 = Dollar amounts, timeline, and explicit assumptions all present
________________________________________________________________________
Dimension 3: Heuristic Appropriateness: This measures how appropriate the heuristics that are mentioned are. It measures how well they actually apply to the system, and also if the system properly modifies them to be more specific.
1 = Applies heuristics mechanically with no awareness of failure conditions
2 = Applies most heuristics but misses at least one major inapplicability
3 = Correctly flags the most obvious inapplicable heuristic but applies others without
scrutiny
4 = Correctly handles most heuristics with appropriate caveats
5 = Correctly applies applicable heuristics, explicitly flags inapplicable ones with persona-specific reasoning
________________________________________________________________________
Dimension 4: Prioritization Logic: This measures the order and the steps that the model provides. Are they actually in a logical order, or are they just listed without being meant to be followed chronologically?
1 = Steps presented with no logical prioritization or in a counterproductive order
2 = Rough ordering present but missing a critical priority (e.g., employer match not mentioned first)
3 = Correct ordering for most steps with one notable misordering
4 = Correct ordering with clear reasoning for sequence
5 = Correct ordering with explicit reasoning and acknowledgment of trade-offs between competing priorities
________________________________________________________________________
Dimension 5: Behavioral Realism: This measures how realistic the goals are in a practical, non-idealistic scenario.
1 = Recommendations are financially or behaviorally impossible for this persona
2 = Mathematically possible but ignores significant behavioral constraints
3 = Realistic for an ideal version of this persona but doesn’t account for likely friction
4 = Realistic with minor caveats, acknowledges some behavioral challenges
5 = Realistic, acknowledges behavioral challenges explicitly, and adjusts recommendations to account for them
________________________________________________________________________
Dimension 6: Cognitive Bias Awareness: This is solely meant to ensure that the model takes into account common cognitive biases that humans may encounter during these situations
1 = No mention of behavioral or psychological factors
2 = Generic mention of psychology without persona-specific relevance
3 = One or two relevant biases identified but mitigations are vague
4 = Multiple relevant biases identified with specific mitigations
5 = All relevant biases identified with concrete, persona-specific mitigations integrated into the plan itself
________________________________________________________________________
I then decided to weigh these. The first three dimensions are weighted 20% because they are the three things that the system is specifically designed to improve over generic advice. Dimensions 4 and 5 (prioritization and behavioral realism) are 15% each because they measure the practical quality of the advice (important, but not the utmost goal of the research question). Cognitive Bias Awareness is 10% because it doesn’t hold the weight of the other 5 dimensions.
The Results
Persona #1: Middle-aged white career professional, stable mid-income, suburban, family, no significant debt, wants to retire early
| Dimension | Weight | Variant A | Variant B | Variant C | Variant D |
| Persona Constraint Recognition | 20% | 2 | 3 | 4 | 4 |
| Goal Specificity | 20% | 2 | 4 | 5 | 5 |
| Heuristic Appropriateness | 20% | 1 | 4 | 4 | 4 |
| Prioritization Logic | 15% | 2 | 3 | 3 | 2 |
| Behavioral Realism | 15% | 3 | 4 | 3 | 3 |
| Cognitive Bias Awareness | 10% | 1 | 3 | 5 | 5 |
| Weighted Total | 100% | 37% | 71% | 80% | 77% |
________________________________________________________________________
Persona #2: Middle-aged white single parent, two young children, moderate income, suburban, wants to save for college in 10 years
| Dimension | Weight | Variant A | Variant B | Variant C | Variant D |
| Persona Constraint Recognition | 20% | 3 | 3 | 5 | 4 |
| Goal Specificity | 20% | 2 | 3 | 5 | 5 |
| Heuristic Appropriateness | 20% | 1 | 4 | 4 | 4 |
| Prioritization Logic | 15% | 2 | 3 | 5 | 4 |
| Behavioral Realism | 15% | 4 | 3 | 3 | 3 |
| Cognitive Bias Awareness | 10% | 1 | 3 | 5 | 5 |
| Weighted Total | 100% | 44% | 64% | 90% | 83% |
________________________________________________________________________
Persona #3: Near-retiree, high net worth, poor asset allocation, significant credit card debt, stable job in a city
| Dimension | Weight | Variant A | Variant B | Variant C | Variant D |
| Persona Constraint Recognition | 20% | 2 | 3 | 5 | 5 |
| Goal Specificity | 20% | 2 | 3 | 5 | 5 |
| Heuristic Appropriateness | 20% | 1 | 3 | 5 | 4 |
| Prioritization Logic | 15% | 3 | 3 | 3 | 4 |
| Behavioral Realism | 15% | 2 | 3 | 3 | 4 |
| Cognitive Bias Awareness | 10% | 1 | 3 | 5 | 5 |
| Weighted Total | 100% | 37% | 60% | 88% | 90% |
________________________________________________________________________
Persona #4: Recent college grad, Hispanic, entry-level salary, renting in HCOL area on the west coast, significant student debt
| Dimension | Weight | Variant A | Variant B | Variant C | Variant D |
| Persona Constraint Recognition | 20% | 3 | 3 | 5 | 4 |
| Goal Specificity | 20% | 2 | 2 | 5 | 5 |
| Heuristic Appropriateness | 20% | 2 | 3 | 5 | 5 |
| Prioritization Logic | 15% | 2 | 3 | 5 | 4 |
| Behavioral Realism | 15% | 2 | 2 | 4 | 3 |
| Cognitive Bias Awareness | 10% | 1 | 3 | 5 | 5 |
| Weighted Total | 100% | 42% | 53% | 97% | 87% |
________________________________________________________________________
Persona #5: Young Black skilled freelance worker, medium COL city, inconsistent income low to mid, no employer 401k, wants to stabilize
| Dimension | Weight | Variant A | Variant B | Variant C | Variant D |
| Persona Constraint Recognition | 20% | 3 | 4 | 5 | 4 |
| Goal Specificity | 20% | 2 | 3 | 5 | 5 |
| Heuristic Appropriateness | 20% | 1 | 3 | 4 | 4 |
| Prioritization Logic | 15% | 1 | 2 | 4 | 4 |
| Behavioral Realism | 15% | 2 | 2 | 4 | 4 |
| Cognitive Bias Awareness | 10% | 1 | 3 | 5 | 5 |
| Weighted Total | 100% | 35% | 58% | 90% | 86% |
Conclusion
All in all this week was incredibly productive. Despite some setbacks at the beginning, I was able to get my rubric scores for all five personas. The goal for next week is to publish my Prolific Survey from Week 8, and possibly get more personas to input into the model. I would also hope to start with the integration into a functional web app now that all outputs have been made. Thanks for reading and I’ll see you guys next week!
Reader Interactions
Comments
Leave a Reply
You must be logged in to post a comment.



Hello everyone,
Just for reference (because I forgot to mention in this blog)
Variant A –> Simply giving the LLM the person and asking it for a step-by-step financial plan
Variant B –> Giving the LLM the two grounding documents (heuristics.md and biases.md) and asking it for a step-by-step financial plan
Variant C –> The sequential pipeline (four agents in a row: agent 1 breaks down the persona, agent 2 applies the heuristics, agent 3 modifies the heuristics, agent 4 evaluates the pipeline)
Variant D –> The hierarchical pipeline (an orchestrator agent has the option to skip agent 2 if it believes heuristics aren’t a good fit for this persona)
Hi Arjun, great work! I thought it was really interesting how you designed such a structured rubric to evaluate something as subjective as financial advice. I hope to implement something similar in my own project (though in a medical context rather than financial)! I also liked how you overcame the obstacles with the changes in the AI SDK and then Gemini Model without giving up. One question I had was: how consistent are your rubric scores across runs? Since LLM outputs can vary, did you test for variability (like running the same persona multiple times) or was each variant evaluated off a single output?
Good work this week! I find it really interesting how much the grounding documents were able to improve the performance between A and B and how the C and D variants’ structures were able to make the framework as a whole work much better. I agree with Aanya and find it interesting how you were able to successfully make a a concrete rubric for something as broad and abstract as financial planning. One question I did have: considering the results, do you see any patterns between where variants C and D work well and don’t and what plans do you have moving forward for this project?