WHOI Summer Fellowship Research
- Researched, analyzed, and created data visualizations out of old ship logbook weather reports
- Wrote a 16-page report on the interdisciplinary tools used in the research and the potential impact of this work
- Presented my work to an audience of 50+ scholars at Woods Hole Oceanographic Institution through a poster session and a talk
Description of the project taken from the final report I wrote at the end of the summer:
“During oceanic expeditions, pre-modern sailors meticulously recorded information on their longitude and latitude, the local wind conditions, and the state of the sea. The nature of the data changed as new measurement instruments and reference points were introduced. For a long time, sailors provided qualitative recordings of wind speed instead of quantitative (e.g.: “light breeze” instead of 5 meters/second). For that reason, this data requires additional processing before being usable for climate modeling.
In particular, the unique phrases used in wind descriptions must be classified into one of thirteen base wind force levels to be accurately translated into a numerical value. Manually categorizing this information takes an incredibly long time. Because of historical weather data’s importance for climate science, we investigated if machine learning could speed up this process while producing accurate results.
Here we show that k-means nearest neighbors clustering, while efficient, generates outputs with reduced accuracy when compared to the data classified by humans. However, there is a noticeable improvement in the quality of the clustering when we introduce the Beaufort Wind Force Scale’s thirteen categories as starting centroids.
These results show that machine learning could be a useful tool for wind term processing and that well-placed human input aids in the accuracy of outcomes. Therefore, we intend to employ cross-validation methods to help with the interpretability of the machine models utilized. Additionally, we plan to test how neural network clustering fares, as it has more room for iteration and human input.”
NLP, Machine Learning, and Linguistic Revitalization for Tupian languages: a review
- Wrote a literature review
- Gave a talk at the MINK-WiC 2023 conference
- Presented at the Grinnell College Computer Science poster session
An in-progress review of the current computational linguistics
literature as it applies about Brazilian indigenous languages,
specifically those from the Tupi family. This includes a survey of
frameworks, tools, and linguistic databases. The review focuses on innovative machine learning models for processing low-resource, morphologically rich languages.
- Co-founded a series of interdisciplinary discussion panels
- Planned, organized, and advertised the first event
- Facilitated a scholarly discussion about technology
As a GrinTECH cabinet member, I co-founded an interdisciplinary series of discussion panels about technology. I facilitate the panels, which centers academics from different departments. This format brings nuance to
discussions about the impacts of technology and aims to foster a
welcoming environment for people from all academic backgrounds.
The first one took place September 21st, on the topic of “Social Networks, Hierarchies, and Change”. We had a great turnout of 20+ people with diverse interests and majors. We talked about the importance of machine learning interpretability for AI fairness; how the usefulness of new technologies depend on the context in which they are implemented; invisible networks and hierarchies, what that means for tech; and more. Attendees appreciated how nuanced and well-facilitated the discussion was. The panelists also enjoyed how well-structured the event was and said it was a great opportunity for them.
I planned the event, spread the word about it, and gave a speech about my personal experience with Joy’s poetry before her reading. Joy loved my introduction and it was so lovely getting to share that moment with her.
- Multiple BRASA (Brazilian Student Association) events on campus
Highlights: the November Fest (a celebratory event with traditional music, food, and games), BRASA’s participation in the Spring 2022 Hometown Expo (I developed and printed pamphlets about special Brazilian Portuguese words), and two dances for Grinnell’s Cultural Evening.
An investigation of how textual analysis algorithms interpret the works of Fernando Pessoa, a Modernist Portuguese writer who wrote as though he were other people, making great use of these heteronyms. My leading question was if the algorithm would be able to tell that these different heteronyms were, in reality, one person through methods of authorship analysis. The result highlighted the stylistic proximity between certain heteronyms, as well as the range of Pessoa’s prose. This was my final project for Professor Erik Simpson’s Spring 2022 “Lighting the Page: Digital Methods in Literary Studies” class.