Showing posts with label Grants and Funding. Show all posts
Showing posts with label Grants and Funding. Show all posts

Friday, October 21, 2016

Big Data and the Mortuary

Sounds gross, right?

But here's the story. I get on the "big data" bandwagon.
I develop these nice scalable algorithms for learning from the big data. I test them on data that does not fit in the memory of my personal laptop (otherwise the reviewers for my journal papers do not agree that I am doing GOOD research). I use a cluster for all of this work - which I have spent hours configuring and playing around with.

But then I decide to relocate. With all my tangible possessions. How do I move this large cluster?

Needless to say, I am emotionally attached to it, by now. It was not easy getting it up and running in the first place. Then I did all the dirty data cleaning work on it. That was another 100+ hours of my existence. And now I am facing this scenario -- I do not have the physical space to store the machines; the administrator who spent many sleepless nights with me has decided to take a break.

The research organizations and foundations that helped set up the cluster have not given thought to the problem. They are only concerned with getting useful things out of it still. Once you build an empire, you usually do not want to relocate.

The environmentalists roll their eyes at me. Large scale computing infrastructure in the electronic waste recycling? You serious? Well go find a place to host them.

Yes, true.

Meanwhile, does any one care about what would happen to these large scale repos once no one rides the bandwagon anymore?

I feel like the Lorax in Dr. Seuss's book(http://www.seussville.com/books/book_detail.php?isbn=9780385372022).

The Onceler and the Truffula trees had their way. The Onceler is unrepentant. He  defiantly tells the Lorax that he will keep on "biggering" his business, but at that moment one of his machines fells the very last of the Truffula trees. Without raw materials, the factory shuts down. The Lorax says nothing. Just one sad backward glance and disappears behind the smoggy clouds. Where he last stood is a small monument engraved with a single word: "UNLESS". The Onceler ponders the message for years, in solitude. 

Tuesday, October 20, 2015

Workshop on Data Science, Learning and Applications to Biomedical and Health Sciences

The NSF sponsored workshop on Data Science, Learning and Applications to Biomedical and Health Sciences (DSLA-BHS2016) will be held on Jan 7-8, 2016 at the New York Academy of Sciences. Workshop website: https://sites.google.com/site/dslabhs2016/
Please contribute if relevant to your research.

Saturday, October 10, 2015

Measuring influence in prospographical research

Suppose you wanted to know the life story of the Australian operatic soprano, Dame Nellie Melba who was famous in the late Victorian era and early 20th century. Where would you begin your search? You could start reading about her life on the internet (perhaps on Wikipedia) and follow-up with references therein. Very soon, however, you could find yourself digging through stacks of old newspapers such as "The Musical Times" or "The Sun" published from New York to learn more about her debut at the Metropolitan Opera. Information is sparse and often incomplete.

Relying on secondary biographical information such as family archives and photographs, publicly available archives including newspaper articles, financial accounts from cities, economic and fiscal sources such as sales of deeds and tax lists, and other surviving documents of the era helps weave a story of the subject's life. It helps verification of facts from multiple sources.

A team of researchers from the School of Management, SUNY Buffalo (Dr. Haimonti Dutta), the Department of Computer Science at IIIT-Delhi (Aayushee Gupta), IBM Research India (Dr. Srikanta Bedathur) and TCS Research India (Dr. Lipika Dey) have been involved in this prospographical research. The project was funded in its initial phases by the National Endowment of Humanities.

Article level data was obtained from the historical newspaper archive of the New York Public Library after curation by the New York Public Library Labs. Using techniques from natural language processing and large scale machine learning, the team was able to build a system that could identify influential people. The noisy text from old historic newspapers were subjected to Optical Character Recognition (OCR) and the text from articles was spell corrected. This was then used to form a people gazetteer from which people with influence were identified. It was particularly interesting to find local people who held a sway in the government offices, arts and sciences, and the armed forces at that time.


A small set of influential people detected using their algorithm from two months of newspaper data published in ``The Sun" newspaper.
The team is now in the process of building a historical timeline very similar to the jazz timeline hosted by pbskids.org to help kids learn about influential people local to a geographic region.

Related Publications
1. Aayushee Gupta, Finding influential people from a historical news repository, 2014. [Master's thesis] 
https://repository.iiitd.edu.in/jspui/handle/123456789/166
2. Aayushee Gupta and Haimonti Dutta, Evaluation of Spell Correction on Noisy OCR Data. INFORMS Workshop on Data Mining and Analytics at INFORMS Annual Meeting, Philadelphia, October 2015.
3. Aayushee Gupta, Haimonti Dutta, Srikanta Bedathur and Lipika Dey. A Machine Learning Framework for Prosopographical Research. In preparation.