15.6 Tier Lists are here! Banners and Cumulative List have been updated accordingly.
Ubers are evaluated based mostly off of stats and niche. Let’s use a theoretical new uber for example. If a new uber were to be a kamikaze unit with similar stats & niche performance to Yukimura, then it’s already pretty clear how good that uber is without much actual in-game testing being needed since Yukimura has already been thoroughly examined.
Pretty much every uber falls under that category. But for new ubers with unique mechanics that can’t be evaluated completely off of stats, testing has to be conducted, in which the uber goes through a selection of stages I think it would be good at and see how it fairs.
These ubers are tested at level 60, with max talents if available. Along with that, most stages will be tested on the grounds of End Game. Ubers and units are also allowed to be tested with their optimal talent orbs, however this excludes dojo orbs.
Most lineups are composed of a full row of non-uber research combos with Rock as the primary meatshield, and either Ramen/ Jiangshit as secondary defense. However, more meatshields and completely different lineup setups are allowed. I am listing this because this is my preferred method of testing most ubers as Research + Rock creates extremely dynamic gameplay for maximum control.
Non-uber units are also accompanied at their max level and talents, being at level 50 for gacha rare/ super rare, +90 for normals, etc.
If an uber is suspected to be top tier, it will be ran through the current hardest stages in the game. This tests the uber on the harshest conditions of the game. The assumption is that if the uber performs excellently on the hardest stages, it will perform just as well or better on easier stages.
What's considered the hardest stages changes frequently with how the non-uber meta evolves. This is because stage difficulty is judged on non-uber difficulty, the reasoning being that PONOS has confirmed that their stage design philosophy revolves around balancing stages without hard to get/ limited units. This clearly includes uber rare cats as this game design logic has stayed consistent throughout the entirety of Battle Cats' lifetime (even with the large quantity of free ubers given to players).
Below is the current list of hardest stages that potential top tier ubers are tested on. Note that this list is not in any specific order.
Feathergeddon: Swan Song (Poultrio stage 4)
Babies First (Doremi Revenge)
Heinous Road (Villainous Gods)
Endless Forms Most Beautiful (Darwin)
Filibuster Invasion (Filibuster Outbreak)
King of the Beasts I (Xenobeast Stage 1)
King of the Beasts III (Xenobeast Stage 3)
Rank-Up Test 3 (Dojo 12-3)
The Belated Priest (3 Crown)
Glass Slippers (3 Crown)
Adventurers Journal (3 Crown)
BCU (Battle Cats Ultimate) is an independent recreation of Battle Cats, describing itself as a fan-made emulator.
Sourcing testing from BCU is valid and is treated as proof of concept, but definitive conclusions based on emulation comes with significant problems.
BCU is independently written primarily in Java (Kotlin in the Android version) whereas the official Battle Cats by PONOS is programmed and ran on an undisclosed game engine.
This fact alone already creates massive uncertainty in how these versions behave compared to each other such as the order in which actions are processed in a frame, decimal precision/ rounding, timings, and random number generation which are likely not exact across official and emulated versions. The existence of these small differences and quirks between the versions are strongly supported with the large quantity of BCU-specific bugs (such as surge bug and documented issues with LD attacks, issues that never occurred in the official app on the original game engine)
There is no documentation of BCU being tested frame-by-frame compared to the app, so these specific discrepancies cannot be proven. But the project history suggests that it is also unreasonable to assume perfect equivalence especially with how BCU is a recreation and not the original game.
The argument against this is that "BCU emulates the official app accurate enough", however this statement depends on the claim. BCU is accurate enough to explore strategies and approximate units, but it can't be used as definitive proof that a lineup or featured unit will win in the official app. To prove that definitively, the test must be done on the official app, as simple as that.
BCU supports custom units, enemies, stages, abilities, attack behavior, LD ranges, base distances, animation speeds, and numerous mechanics beyond ordinary conditions. Its documentation describes a highly customizable experience.
The program has also supported controls involving unit levels, costs, custom combos, custom talents, custom traits, custom abilities, and other battle altering options.
This does not mean every BCU user is maliciously altering their runs. It means that a victory screenshot by itself cannot verify the exact BCU build, the emulated game data version, RNG values, speed settings, or TAS tools. Even an honest player may unknowingly be using a configuration that differs from the official app.
Of course there is the possibility of intent to falsify runs in the name of evidence, something that is much more difficult to evade with recorded runs in the official app.
Certain mechanics and behaviors in BCU have had very stark parities between the official version that have been corrected. This is a direct example of game behavior differing between the two versions.
BCU features many bugs that are continuously being fixed. Here were some of the most notable examples, some of these bugs may or may not have already been fixed:
There existed a bug related to surges that allowed them to hit an extra time, giving the surge user higher damage and longer procs. This causes units like Uril to be dramatically stronger on stages like Villainous Gods where he can perma-slow Aku by himself.
The sniper cat powerup has finicky properties. Sniper's projectile had the possibility of hovering in the air and missing completely.
At one point conjures couldn't hit an animated enemy base if it was not flagged as a boss.
Knockback boundary was set extremely far from the enemy base for bosses at some point. This caused the Li'l Lion + iCat cheese on Poultrio stage 4 to change dramatically in making it easier to perma-freeze but likelier to clip past the boss and lose the run.
Sometimes Zombie corpses would slide forward.
Sometimes death surges played their animation but did not spawn a surge.
Sometimes explosion units and enemies would create an explosion even if their attack completely missed.
These reports do not prove that every current unit is inaccurate. Regardless of whether these bugs are fixed or not, what this proves is that BCU parity has been historically incomplete and that there can be differences in actual battle compared to real gameplay.
BCU features mechanics such as greenbox (Cat CPU for individual slots), speed up beyond 3x, and TAS. These tools do not exist in the official Battle Cats while providing advantages for the BCU user which is why any run that utilizes these tools constitutes as cheating and are disregarded entirely.
When testing ubers, there will be occasions where an uber is strong enough to the point where you only need a few slots or less just to be completely hard carried by the uber on a stage. An example of this is Luna who with Corn Cat and Uril can defeat the entirety of Dojo 12-3.
The idea is that if an uber is proven to beat a relevant stage with less slots required than another uber, this provides an extremely strong basis for the former uber to be ranked higher as carrying with less slots means a few things:
The uber's kit is more capable, likely filling in multiple roles in battle or being greater at sustaining itself
The uber's dominance carries greater margins, as in practical runs you can use those empty slots to win better/ more easily
This method of unit analysis is valuable information, but it alone doesn't stand as definitive ranking evidence.
The actual objective of a Battle Cats stage is to defeat the enemy base while preserving your own. It is not to leave as many empty slots as possible.
Clearing stages with the least amount of slots is a self-imposed challenge that does not provide any benefit for a normal player. When a unit performs well under an invented restriction, it does not automatically prove that the unit is more valuable in a normal 10 slot lineup that the game allows you to create.
One central idea of minimum slot testing is that if an uber is proven to beat a stage with very few slots, the average player can copy the strategy and add more units in to decrease the margin of error and attain victory easily.
This line of thought is reasonable at first, but it doesn't hold up well with issues you'll see further below. The shorthand is that it's unreasonable to assume that elements such as positioning, money management, RNG, and specific timings are immediately solved by adding more units into a lineup nor could you assume that victory is easily attainable if the player does not inherit the tester's execution.
Possibility also does not measure the size of margin, a minimum slot clear does not provide information on how "hard" it won the stage aka victory threshold which is another element not instantly solvable by adding more units in a lineup, that assumption would have to be proven with in-game testing.
The game gives you 10 slots, use those slots.
Minimum slot testing prefers units that can handle multiple roles at once. This naturally disadvantages units that have a specialized function to support the whole lineup.
For example, a powerful support unit may dramatically improve multiple units (or the whole lineup) while being unable to beat the stage with a team composed of 3 or 4 slots. Under minimum slot ideology, that unit appears weak because the test removes the environment in which it is designed to excel in.
Battle Cats gives the player 10 slots because lineup construction is one of it's biggest strategy components. Units are supposed to complement each other and many units are built specifically with this in mind. A unit that creates strong results through synergy is not automatically worse than one that functions independently. Consider two hypothetical units:
Unit A can barely clear with a few slots after extremely precise timing and RNG setups.
Unit B cannot clear with a few slots but makes an ordinary 10 slot lineup dramatically safer, faster and more consistent.
Minimum slot ideology declares A superior, even though B may be more useful to nearly every actual player. It confuses independence with value.
A screenshot of a completed stage with minimal slots does not give information on whether it was cleared on attempt 1, once in a hundred attempts, or after extreme RNG favorability.
The relevant statistic should not be “Can it happen?”, it should also be "How reliably can it happen under repeatable conditions?" A lineup with a 90% clear rate is generally stronger evidence than a lineup that succeeded once after hours of resets even if the latter used one or several fewer slots.
Removing supporting units reduces redundancy. That makes the outcome more dependent on individual RNG events such as savage blows, dodge activations, debuff procs, and enemy spawn rates/ timing.
A selected victory may represent the most favorable result out of many attempts. Unless the tester reports every attempt, the evidence suffers from survivorship bias: you see only one successful screenshot but not the potentially hundreds of failed runs that preceded it. On the opposite end of the spectrum, a singular incredibly lucky run may be nearly impossible to replicate.
At extreme low slot clears, the deciding factor may be the tester’s ability to conduct exact unit spawn timings, cannon timings, money management, positioning, or restarting until favorable conditions occur.
That can be impressive as a challenge, but it measures the player’s execution rather than the unit’s performance. Minimum slot clears attribute the entire result to the unit and ignores the strategy itself.
If the run only works inside a narrow timing window, the victory demonstrates that a skilled player can execute that timing. It does not demonstrate that the unit will perform exceptionally for an average player. A unit that has a high skill ceiling under near-perfect timing may still be less practically useful than a slightly weaker unit that succeeds with great margin of error.
This is supported by the previous point of RNG amplification. Suppose a strategy has only a 5% success rate. A determined tester can restart until the strategy eventually works. The victory then looks identical to a strategy with a 95% success rate.
Without knowing this critical information, the two runs are indistinguishable.
"What about units like Balrog and Lasvoss? Aren't these reset heavy/ RNG baked units that require immense effort to achieve victories with? These units ranking high contradicts these statements."
The reason why high skill ceiling units like Balrog and Lasvoss don't apply to these policies is how their potential is easy to replicate under repeatable conditions. With full lineups, setups for victory are reliable to the point where it's executable by the average player for practical usage.
These units also do not take hours of attempts to work for most stages. If they did, their credit would be ruined under the basis of practicality. These units are feasible to use and have repeated, community-wide accounts to suggest consistent net positive results. These are also the exact type of units that benefit from the support of a full lineup and are disadvantaged from minimum slot testing.
This point only applies to minimum slot runs conducted in BCU.
A comfortable 10 slot strategy will remain successful even if targeting or knockback timing differs slightly between BCU and the official app. A 4 unit strategy that wins with the base at low health is likely to fail in the official app because of a single altered interaction between each game engine.
Minimum slot runs intentionally remove safety margins. Therefore, they are the least appropriate type of runs to confirm between an emulator and the official app.
If the BCU run only works through extremely optimal spacing or timing, an inability to reproduce it officially is not a minor inconvenience. It undermines the claim that the BCU result proves an official in-game result.
In real life, policy makers employ a precautionary principle when there is a possibility of harm when making a decision and conclusive evidence is not yet available.
BCU runs are welcome as suggestions for strategy and proof of concept, but they are not definitive evidence of performance for units compared to the app.
Tests that are conducted in BCU will be treated as unverified until reproduced in the official app.
Any ranking claim based on a highly optimized or minimum slot BCU clear must be reproduced officially in-game before it is treated as confirmed evidence.
Minimum slot clears will be recognized as role-compression evidence and are still welcome to conduct for gauging unit performance, but it is not an objective or definitive measurement of overall unit strength.