Every Java developer with more than a few months of coding experience has written code like this before:
What I realized just recently is that Java 7 already provided a fix for this ugly code, which not many people have adopted:
Yay! No Exception! But it is not only nicer, it is also faster! You will be surprised to see by how much!
Let us first look at the implementations for both getBytes() calls:
Not exciting. We shall dig on:
and
Wooha. Well it looks like the one taking a Charset is more complicated, right? Wrong. The last line of encode(String charsetName, char[] ca, int off, int len) is se.encode(ca, off, len), and the source of that looks mostly like the source of encode(Charset cs, char[] ca, int off, int len). Very much simplified, this makes the whole code from encode(String charsetName, char[] ca, int off, int len) basically just overhead.
Worth noting is the line Charset cs = lookupCharset(csn); which in the end will do this:
Wooha again. Thats quite impressive code. Also note the comment // We expect most programs to use one Charset repeatedly.. Well thats not exactly true. We need to use charsets when we have more than one and need to convert between them. But yes, for most internal usage this will be true.
Equipped with this knowledge, I can easily write a JMH benchmark that will nicely show the performance difference between these two String.getBytes() calls.
The benchmark can be found in this gist . On my machine it produces this result:
Benchmark Mean Mean error Units
preJava7CharsetLookup 3956.537 144.562 ops/ms
postJava7CharsetLookup 7138.064 179.101 ops/ms
The whole result can be found in the gist, or better: obtained by running the benchmark yourself.
But the numbers already speak for themselves: By using the StandardCharsets, you not only do not need to catch a pointless exception, but also almost double the performance of the code 🙂
Blog author
Fabian Lange
Do you still have questions? Just send me a message.
Do you still have questions? Just send me a message.