How about google maps says "keep north"..as if I am sitting in my car with a magnetic compass...gets my goat everytime
HN user
mohi13
I am building an AI generated podcast..which creates a fun, informative podcast on news stories impacting the Indian stock market.
One podcast a day
Here it is: https://www.youtube.com/@FoliyoAI
The word is “False Positive”
not working..showing almost empty page
Can also you Dataturks do a similar thing, hooking up with Mturks/your own team..also allows overlapping selection and NER+Classification etc. Adding extra data and much more.
Here is an online demo:
https://dataturks.com/projects/Dataturks/Demo%20Document%20A...
I get disgusted by thinking why does Netflix want people to not sleep. Why auto-play in 5 sec? Isn't this morally wrong, like dangling cocaine in front of a drug addict.
I agree that people have free will but designing products so that people waste days just watching TV is very very sad IMHO.
Here is a comparison we did between APIs from Microsoft Vs Amazon Vs Kairos for face detection. Kairos came out to be much better overall.
"with almost half quitting after the first week"...to me this sounds super critical...no amount of funding or a partnership with even god him/herself, can't make it work unless this is fixed.
thanks, will see how can I block crawl on this dataset. BTW how does it hurt exactly, couldn't find much except in case of safe-search mode.
Would be really interested to see the results of this.
Also, consider using Dataturks to create and host the dataset.
:)..There are a few examples of males as well.
BTW we had a question in general to anyone who can help us, Does hosting such a dataset cause issues with SEO etc? Anything else we should be aware of?
Makes a lot of sense, actually its really difficult to get a large enough dataset for moderation tasks to make a decent inhouse model for a fair enough comparison.
Sure, we can try scraping that from pornhub etc but fee then the negative classes would be very domain specific, using stock images may not provide a good measure.
Also, its really weird to assign such a task to any of your employees, feels kinda strange :)
Here are 1000s of more open datasets for anyone to explore, use or build upon: https://dataturks.com/projects/trending
Funny things, the exchange I had my money in, went bust and I lost most of my Bitcoins..Irony !!
This is just a demo of an attempt to use ML to do price prediction, would be surprised if this would be 100% accurate predictor.
A dataset containing actual search queries on bestBuy.com manually labeled by human experts.
Key Features
~1700 labeled named entity pairs
7 Categories
Human labeled dataset
Thanks for the feedback, but this seems pretty opposite advice to what I have seen some ppl saying, "to ask what ppl want and then build that"?
Yes that's what we are actually trying to fix. What we have seen from personal experience and from talking to people in other places, almost everywhere they build a hacky solution and takes up a full time dev's time to build it and maintain it.
We realised we could build a good solution around this and do a really good job at that so that others can simply use us.
Its similar to github as in you collaborate with your team on a central repo w.r.t datasets.
We are currently looking at feedback from folks on what features we can add more ?
These are the use cases we are actively building out. Should be there in a few weeks.
I invite you to give it a try and provide feedback on how can this be useful for you/your team?
We make it super easy to do ML data annotations. We are online platform for teams to collaboratively build ML datasets.
Using this you can enlist your team/colleagues to help out in annotations, we provide good data visualization to help get insights from your data and tracking on which of the team members if helping the most.
Eagerly looking forward to HN communities feedback. :)
Yeah even we think of it as your personal MTurk. You can do the annotation with your team/network and not have to depend on crowdsourcing it.
Don't you think in that context the name is kinda appropriate?
Hi thanks for the comment and thanks for SpaCy, have used it a few times and found it really powerful.
Also what you are doing with prodi.gy is also pretty interesting with active learning and stuff.
W.r.t privacy policy the data is owned by the users and he can delete it as he sees fit. We will never access it.
Yeah that's true, the above commenter syllogism is the founder of Spacy I guess. Similar tool but for a different purpose.
That certainly is one use case. Even for an auto-generated training data it is almost always the case to have some noise in the data and taking out the golden set from that is rarely an option, we always need to do some manual tagging and cleaning. Thanks for the feedback.
Thanks. We had looked at what they were doing, similar as a concept but quite different in use cases. CE is mostly to support multiple platforms like Magneto/Shopify or other business integrations etc with just one integration, we are making it for developers and teams to integrate and iterate faster..similar to what Segment does for Analytics.
Thanks for the feedback, I guess that makes more sense.
We kind of think it like a Netflix model where you pay X$ per month and can use any API in the package (unfortunately not unlimited usage)
Just curious, how would you describe this? May be would help me to put in a better way.
Thanks for the kind words, Yes in a way this is comparable to segment.
Shoots us a mail for contact@dataturks.com, we could arrange a free plan for our beta users :)