Full disclosure: I'm an employee
HN user
dwynings
I am Dru Wynings.
http://druwynings.com
Feel free to email me about anything: dru@druwynings.com
Head of Marketing @ Sensible.so. Previously Head of Marketing and Business Development at Diffbot (http://www.diffbot.com/), BizDev @ Heyzap (YC W09). Self-taught programmer & designer.
@druwynings
Sensible | Technical Product Marketing Manager | Remote | Full-time | https://www.sensible.so/
Sensible is an API-first document processing platform that streamlines data extraction for developers and product teams.
We're looking for an experienced Product Marketer to join our growing team. As a product marketer, you will lead core product marketing initiatives, build awareness with our target market, and generate inbound leads for our sales team.
This role requires a mix of creative and quantitative thinking. You'll develop a product messaging and positioning framework that resonates with developers and presents our products in ways that increase Sensible’s ACV.
Ultimately, you will drive awareness, product adoption and revenue growth. This position reports to the Head of Marketing.
Sensible is a remote-first company and this role is open to anyone located in North or South America. We provide competitive compensation and meaningful equity.
More info: https://sensiblehq.notion.site/Technical-Product-Marketing-M...
-- If this sounds interesting, please reach out to dru@sensible.so
Not all that surprising given: https://sfstandard.com/2023/10/25/carta-san-francisco-lawsui...
Yeah, we tried to have some fun with it.
It was actually because I didn't want to link to a hubspot url and webflow only supports uploads up to 10mb.
Also available as a PDF if you're interested: https://www.sensible.so/history-of-the-pdf-pdf
Sensible | Senior Infrastructure Engineer | Full-time | REMOTE |
Sensible connects software to the messy, unstructured data that businesses encounter every day. Our goal is to bring about a world where computers do the work that computers are best at (processing large volumes of data) so that humans can do what humans are best at (critical thinking and empathy).
The first piece of this problem that we're tackling is document parsing. Even as software is eating the world, so many business workflows still rely on two entities sending PDFs to each other. In a single afternoon with Sensible, developers can ship a production-ready API endpoint that turns documents into useful data.
Sensible was founded in 2020 and is backed by some of the top investors in Silicon Valley. https://www.sensible.so/about
--
As a senior infrastructure engineer, you’ll work closely with our head of engineering and the engineering team to improve and expand our AWS-based infrastructure, working with technologies like Lambda, DynamoDB, S3, and IAM, to support our production APIs and web app. Outside of AWS we integrate with Google Cloud, Microsoft Azure, and OpenAI. We have a strong testing culture to support our overall reliability, as well as SLAs for our enterprise customers.
--
You might be a fit if...
- You have 5+ years of experience building infrastructure and APIs in AWS.
- You’re product-minded and customer-focused.
- You have experience working in organizations compliant with SOC 2 and HIPAA.
- You are an excellent verbal and written communicator, able to build relationships with different kinds of people across different levels of the organization.
--
Interested? Reach out to the founder josh@sensible.so and mention HN.
--
- Senior Infrastructure Engineer: https://www.notion.so/sensiblehq/Senior-Infrastructure-Engin...
- Customer Success Engineer: https://www.notion.so/sensiblehq/Customer-Success-Engineer-8...
- Product Manager: https://www.notion.so/sensiblehq/Product-Manager-a0dc92f1a9d...
- Account Executive: https://www.notion.so/sensiblehq/Account-Executive-14205252e...
We just closed a lease on a new space where every employee will have their own office. Definitely atypical in Silicon Valley, which meant that buildings with a lot of build out are less desirable. That made it easier to negotiate the price.
We tried Neo4j, but it couldn't support the throughput we needed for injecting facts.
My most common use-case is SaaS trials that require a credit card upfront – I set the limit at $1, so I don't have to worry about cancelling my trial before they autocharge me.
Hey Tegan,
Answered on the parent, but it's somewhat similar.
Fair warning: I work at Diffbot.
Essentially that's what Diffbot (https://www.diffbot.com/) does, except we don't the render pages as an image nor do OCR.
Diffbot renders the page in a headless browser, and uses computer vision to automatically identify the key page attributes and extract normalized data for specific page types (Articles, Products, Discussions, Profiles, Images, and Videos).
This approach enables us to work in any language and on sites that we've never come across before automatically with better than human level accuracy.
At a very high level, it's similar. We use computer vision and ML to extract structured data from any web page, even ones we haven't seen before. https://www.diffbot.com/
If anyone has any questions or wants to try it out, feel free to email me directly at dru@diffbot.com
Disclaimer: I work at Diffbot
Major differences I can see (OP feel free to correct if I'm wrong):
Link.fish
* doesn't provide a web crawler
* relies heavily on microdata, schema.org, RDFa, etc
* relies on manual parsers for sites that don't have microdata embedded
* doesn't full-render pages by default (Diffbot renders every page, so it can use computer vision to automatically extract the data)
* doesn't support proxies
* doesn't support entity tagging
Probably plenty more, but that's what jumps out to me at first blush.
--
Since I see other people have mentioned price as a concern, we're always willing to help out bootstrapped startups. Just shoot me an email: dru@diffbot.com
How do you guys compare yourselves to https://www.yourmechanic.com/, also a YC startup?
Clozemaster looks pretty cool.
A couple notes:
"the post-Duolingo learn language in context app"
This was really hard for me to understand. When I skimmed this, my first assumption was that this was a direct replacement for Duolingo, when it's actually complementary.It would probably be better to split up the long list of languages on to 2 different pages. 1. Language you want to learn, and 2. From which language. An alternative would be to have a separate section for each language to learn on the page with space in between.
The only thing close that I can think of is: https://github.com/sirensolutions/kibi
Demo video: https://www.youtube.com/watch?v=g0O8UNM0B7Y
Honestly I've gotten so much value from Font Awesome over the years that it was great to finally support the project – I have a feeling others felt the same way.
Any ETA on when backers will get private repo access? Super excited for this!
Diffbot (http://www.diffbot.com) | Palo Alto, CA | ONSITE, VISA
We're an AI startup that applies machine learning, computer vision, and NLP techniques to the problem of understanding webpages. Our APIs convert billions of webpages automatically into structured data for the likes of DuckDuckGo, Salesforce, Hubspot, Amazon, Bing, eBay, Adobe and others.
We recently announced our profitability(!!) and raised a $10M Series A by Tencent Ventures and Felicis Ventures.
Looking for ML/CV/NLP specialists, data fusion / knowledge graph, and/or web-scale crawling experts with a track record of building intelligent systems that perform at human-level accuracy rates.
AI Researcher: https://careers.jobscore.com/careers/diffbot/jobs/ai-researc...
Data Operations Product Manager: https://careers.jobscore.com/careers/diffbot/jobs/data-opera...
Search Engineer: https://careers.jobscore.com/careers/diffbot/jobs/search-eng...
Technical Account Exec: https://careers.jobscore.com/careers/diffbot/jobs/technical-...
If you have any questions, feel free to email me directly: dru@diffbot.com
Sure: http://www.diffbot.com/
Sorry about that! If you let me know the result id, I can take a look at what happened.
Yes, that's something we definitely support – you'll want to take a look at our crawler: http://www.diffbot.com/products/crawlbot/
If you'd like me to set you up with a demo, feel free to email me: dru@diffbot.com
Hi Max,
All of our automatic APIs aren't meant to be used on directories or listings pages (for that, we provide a crawler which gathers those links and then feeds them into our automatic APIs). The discussion API is intended to be used on discussion pages, for instance: http://www.diffbot.com/testdrive/?api=discussion&url=https%3...
If you wanted just a list of submissions and score, you could use our custom API toolkit(similar to Kimono): http://www.diffbot.com/products/custom/
Dru from Diffbot here.
Sorry to hear about the issues you ran into when trying out Diffbot! If you have some examples, I'd like to take a look into whether we can improve.
In cases where our automatic extraction using computer vision isn't 100% accurate, we offer a visual interface for actually overriding our default extraction. This input is then used as additional training data for the ML models.
Mozenda et al. are great if you only need data from a couple of sites and you don't mind spending time manually specifying and maintaining CSS selectors for each website and page layout you need data from.
Our crawling, and proxy support, is fairly robust thanks to our hiring the creator of Gigablast [https://gigaom.com/2013/09/10/diffbot-brings-big-time-search...].
If you'd like to give Diffbot another go or you have some examples where the extraction could be improved, please let me know!
Diffbot • http://www.diffbot.com/ • Palo Alto, CA • REMOTE • VISA
We're an AI startup that applies deep learning, computer vision, and natural language processing techniques to the problem of understanding webpages. Our APIs convert billions of webpages automatically into structured data for the likes of DuckDuckGo, Bing, Digg, Instapaper, eBay, Adobe and others.
We use very few 3rd party frameworks and strive to develop our own performant machine learning techniques.
Looking for published ML/CV/NLP specialists, data fusion / knowledge graph, and/or web-scale crawling experts with a track record of building intelligent systems that perform at human-level accuracy rates. If interested, send us a note at jobs@diffbot.com (due to limited HR bandwidth, naked resumes will be discarded).
If you have any questions, feel free to email me directly: dru@diffbot.com