This is really useful, thanks for sharing.
Have you actually manually mapped the similar characters for the entire unicode set? It would be cool to have some offline image processing done that identifies related characters to augment what you already have.