Our GCP is down
HN user
mlb_hn
playing with little space aliens at julius.ai
nice overview of progress over time. are there quant metrics for the sim capabilities or is it mostly vibes?
he was good people
yeah, that's a very big caveat - haven't checked neo 20b yet. I've had a hard time getting the AI21 models to use it and those are also pretty big so it's interesting why sometimes it works and sometimes it doesn't. (and Davinci > Codegen Davinci > Curie > J-6B). Fine tunes can also learn to do the inner monologue as well which is really cool - not sure how much is architecture vs. training parameters.
definitely. also works on text translation/comprehension like emojis! https://aidungeon.medium.com/introducing-ai-dungeon-translat.... For actual benchmarks, scratchpad improves GPT-Davinci WIC from 50% accuracy (chance) to nearly 70%.
I think the check and validate is a different sort of scratchpad but maybe not. Seems like at least 3 types - soe for pulling implicit info out of the network viz wic, sometimes for intermediary steps viz coding, sometimes for verification like here.
I get the tokenization argument and it may influence it a bit, but I suspect the n-digit math issue has to do more with search the way it samples (in the bpe link gwern references some experiements I'd done with improving n-digit math by chunking using commas, http://gptprompts.wikidot.com/logic:math). I think since it samples left to right on the first pass, it's not able to predict well if things carry from right to left.
I think can mitigate the search issue a bit if you have the prompt double-check itself after the fact (e.g. https://towardsdatascience.com/1-1-3-wait-no-1-1-2-how-to-ha...). Works different depending on the size of the model tho.
Couple things there where you can see if it improves with the prompt/formatting. E.g. with Davinci (and J a bit but didn't test too much) you can get bette results by:
- Using few-shot examples of similar length to the targets (e.g. 10 digit math, use 10 digit few shots)
- Chunking numbers with commas
- Having it double check itself
and here it's not doing any of those things.An alternative of that is as Oppenheimer once said, sometimes things are secret because a man doesn't like to know what he's up to if he can avoid it
This has been a systemic issue reported on for years; e.g. reported by the Intercept in 2017 [1] and Atlantic in 2019 [2]. Not really made clear from the story considering the Economist headline is almost identical to the Atlantic one.
[1]https://theintercept.com/2017/11/02/war-crimes-youtube-faceb...
[2]https://www.theatlantic.com/ideas/archive/2019/05/facebook-a...
I think that goes back to a Karpathy quote [1], don't know where he got it from
My thinking there wasn't because of BPEs, I think it's a graph traversal issue.
Also, better prompt design if you have make implicit meaning explicit can improve the WiC score (http://gptprompts.wikidot.com/linguistics:word-in-context) and ANLI score (http://gptprompts.wikidot.com/linguistics:anli).
Just priming an immediate availability response is likely going to get poor results.
On the other hand, this does bring up an important point, which is that few people have been systemically trying to figure out how to get it to reason through problems. For instance, if you try the pure completion on WiC you get 50% chance (like in the paper) but if you improve the prompt to self-context stuff you raise it to almost 70% (http://gptprompts.wikidot.com/linguistics:word-in-context).
Great article. It's worth considering the Japanese reaction to the bomb - people thought it was cluster munitions or other conventional explosives leading to them taking sub-optimal follow-on steps. While the physicists figured it out, it was not obvious at the time to many people who experienced it firsthand what happened.
I don't think it's going to be about a single prompt; reverse engineering multiple prompts interacting with themselves is hard. There's a lot of cool things to be done with:
(a) creating a pipeline of prompts that combine outputs of previous prompts into new prompts in a predefined manner
and (b) designing prompts to generate other prompts
I think it would be jumping to conclusions to say because the military has vulnerabilities that the entire military industrial complex is a fraud. Pieces of it certainly are sub-optimal and arguably fraudulent though such as DCGS (the "opportunity to improve" terminology used for it in DOD speak means failure)
That's the dataset used in the paper.
Leadership in the US Government wasted a lot of time trying to play down the pandemic [1] as did large (conservative) news organizations [2]. Both are now trying to blame China, arguably to shift blame from themselves.
It didn't need to be that way; South Korea's government action back in January helped them while the US Government's action hindered response [3].
[1] https://www.nytimes.com/2020/03/15/opinion/trump-coronavirus...
[2] https://www.washingtonpost.com/lifestyle/media/on-fox-news-s...
[3]https://www.reuters.com/article/us-health-coronavirus-testin...
I have a WSJ subscription. I don't know of any easy way around their paywall, makes it one of the special cases for scraping because you would actually need to login.
I'm not sure I would agree that the article doesn't matter that much. It's an argument that while the FCC's going along with what the courts decided, they're not necessarily doing so in good faith because they tried to bury the announcement with fluff and didn't use a title that would be understandable to the average user.
What do you mean by dictate what you see? Aside from the sponsored content/ads, isn't it just a prioritized queue of what your Facebook friends posted?
Is he wrong?
Like a telco, Facebook allows people to communicate with one another without going through another intermediary. However, it's not one-to-one like your phone, it's one-to-many.
Like a newspaper the messages are broadcast out to everyone that's subscribed to receive them. However, there's limited editorial control by Facebook (except for content breaching their policies).
If we thought of Facebook as a glorified listserv with a pretty UI (and a bunch of tracking cookies and injected ads) would it still seem like it should be treated as a publisher?
Here's Facebook's business help link for how to upload and use point of sale and other offline data: https://www.facebook.com/business/help/1142103235885551?id=5...
Masquerading as Red Cross can be a grave breach of Article 37 of the 1977 Additional Protocol I [0]
[0] https://ihl-databases.icrc.org/ihl/WebART/470-750111 (paragraph 3.f)
Is there a legal basis for that argument or is that your opinion?
Espionage/sabotage is not a war crime https://ihl-databases.icrc.org/customary-ihl/eng/docs/v2_rul...
The money's one thing, but my concern is that the same ad-fraud tech used to fake all those views can fake views to actual news sites (and influence coverage via analytics). If it can beat Google and Amazon's fraud detection, how well would the news sites be able to detect that sort of attack?
I haven't had any luck getting BBC's numbers on their click-fraud detection rates ='( 3 FOI requests rejected due to security exceptions.
Yep, there's been some academic work like MapWatch (2016) [0] on tracking the differences in how different providers present different borders in different regions over time.
> Sad to see once decent, thoughtful operations like WSJ, NY Times, WaPo, basically turned into conflict generation drivel producers.
I'm not sure that's a fair description of news orgs. The world's fairly complex and you're going to end up with slant one way or another however you try to describe it.
It seems to conflate search/auto-complete and then ignores the context of search having adversarial groups consistently trying to manipulate search rankings. While it's possible this was a good faith article, given News Corps' broader conflicts with Google recently I'd guess it's intentional.