So, make of this what you will, but I do research on psychological measurement.
Also, I have no idea what Goldman is planning to use, so there's that.
However, all of the arguments you're making about bias, etc. could be leveled against everything that goes into a hiring decision--absolutely everything. Education, interviewing style, skills with regard to social norms in a certain setting, everything. Notice, for example, in the article that Goldman has started to consider people from non-ivies, as if it should be some revelation that there are competent individuals from non-ivy schools. It's hypocritical to be argue about personality test bias when the bias due to those other procedures dwarfs that.
The advantage of personality tests is that they're standardized, so you can quantify the bias. This is routine. You can argue about how to quantify it, but there are solid methods for doing so, and it's better than just going with your gut.
As for people mentioning previous use of personality tests: the Myers-Briggs does suck. It's why no one studying individual differences scientifically has paid attention to it for decades. If Goldman comes out using the Myers-Briggs, then complain about that. But complaining about personality tests "because Myers-Briggs" is like saying the internet shouldn't be trusted because "ActiveX on Windows 95."
About reliability: yes, there's issues about reliability, but the reliability of interviews is even worse. There are some people who make solid arguments that you shouldn't even do an interview because they have such little validity in predicting anything above and beyond what's discernable from resumes, history, and test scores.
The one argument you might make about these tests that has some weight is their "fakability." This is their Achilles heel admittedly, although the issue is more complex than it seems because people aren't as good at faking as they think they are, and it's unclear the faking is any easier than on an interview.
One reason why you're seeing increased interest in personality tests now is because of methods that have been developed for counteracting faking. There's been a lot of movement in this area over the last 10-15 years or so, and there's more ways of handling it than there used to. So, for example, in addition to quantifying a faking "style", which you could do before but was of unclear utility, now you can administer items that are controlled for social desirability but measure other dimensions (for example, have people choose between three scenarios that have been shown to be of equal apparent desirability but differ in their level of some other trait).
To be honest, expect a lot more of this. Big IO and testing corporations have been putting a crapton of money into this area and a lot of companies are showing a lot of interest in it, as they're complaining that intellectual ability and "job skills" (e.g., knowledge of algorithms) isn't really the main issues they have with employee performance--the problem is this other stuff, like emotional stability, social skills, etc.