I've been wondering about the distinct lack of incentives for academics to maintain the code and software they develop. (As some people have noted, by the way, the situation is getting better, but it still feels like a crapshoot.) One type of incentive that's not usually available for software maintenance is taxpayer or nonprofit money. Sometimes the (very generous) Sloan foundation will swoop in and unexpectedly bestow millions of dollars upon a deserving project, but that's the exception, not the rule. And altruism doesn't pay the bills. I wonder if the solution isn't fairly simple: quit giving stuff away for free, and make the users pay for what they use. Academics are freeloaders, and they should start bearing some of the cost for the tools they use.
Don't get me wrong: I'm not arguing that software should not be open source (it should), nor am I arguing that it shouldn't be free to download and use (it should, in most cases.) But beyond that, if a community wants to prevent code from rotting, they should bear the cost of maintaining it.
My suggestion is this: once I've released software, I'm not obligated to keep spending my own time maintaining it unless I personally want to or can derive some other type of incentive from my efforts - and as I've already discussed, those incentives are in short supply. So I'll continue to develop software and milk it for as many publications/conference presentations as it's worth. Then I'll move on to other projects.
Beyond that, it's up to you: every time you donate my currently hourly rate to a project, assuming I can make time, you've bought me for one hour to work on fixing bugs, developing new enhancements, or improving documentation for that specific project. (And if I can't make time, at some point I'll be able to hire someone who can.) If people are interested in my open source projects and don't want them to rot, they can either contribute their own time and effort to submit changes, or they can donate the monetary cost of getting me to do it. Fair enough, right?
This provides a decent way to gauge whether working on a particular project is actually a productive use of my time. It also seems like this is something that could be directly mentioned on my CV: I developed Open Source Project X, which has raised $Y in community funding, etc. Creating an open source community is a fuzzy concept, but dollar amounts are something that everyone can understand.
I don't personally expect the floodgates to open at this moment and the donations to start pouring in for anything I've built, but I'm looking to the future. I was raised as an undergraduate in several research groups that produced software and were entirely dependent on federal money. I plan to continue doing software development, and need to think about sustainability.
In 2010 I created a prototype of a programming language called Scotch (the dynamic, functional lovechild of Python and Haskell) largely as an educational exercise while I worked through Krishnamurthi's Programming Languages book. After a few blog posts took off on Reddit/Hacker News, it seemed like there was a decent amount of interest. (Matz, the creator of Ruby, even tweeted about it! Major nerd moment for me. Oh, and there's a mention of Scotch on the Haskell Wikipedia page, just waiting for some very justified moderator to remove it.) It's a pretty complex language and would've required a lot of work to fully develop myself. Reality set in - after the birth of my son, while working on other projects and taking 18 credits of courses, I just didn't have time to freely donate to this project anymore, and while people seemed to think it was cool, no one was helping me to write it.
I created a Pledgie campaign so that people could support the continued development of the language, and, like any good street performer, seeded it with $50 to make it look like people were donating. In the years since I opened the Pledgie, some anonymous saint has generously donated $1.
No happy ending to this story: at some point, I broke the interpreter, and I really don't have time to get back into the code base and figure out where I went wrong. The lack of actual support for the project reveals the true level of value people place on the project - at best, it's of niche interest. I'll probably continue to work on it every now and then, but it's a low priority. I'm not going to invest too much time on it. I have other more valuable ways to spend my time that are more likely to result in direct short-term benefits.
Are there any successful examples of projects that are funded by community contributions instead of grants in academia?
Showing posts with label rants. Show all posts
Showing posts with label rants. Show all posts
Sunday, February 17, 2013
Saturday, January 12, 2013
Academic journals are an obsolete historical appendage of academia
I don't know if it's in my best interest to write this, but it's been a sad day and I feel like it needs to be said, again and again, by everyone who's been affected.
Historically, when information was spread exclusively via ink on paper, academic journals provided a crucial service to the academic community: distribution.
Over the past few decades, the advent of the internet has fundamentally changed the way we use and share information, but traditional journals remain a staple of academia - not because they continue to contribute unique value, but because they're entrenched in the system and careers still depend on publishing in them.
The internet makes publication of information easy and efficient. Modern academic journals don't provide any additional value that could not be easily and cheaply replicated, and because of their history as a physical medium they have been slow to realize the full potential of the publication of information via the internet. Journals still publish discrete issues, enforce page limits, and sometimes charge to print color figures. In the digital age, all of these practices are unnecessary - laughably so. They also charge exorbitant amounts for access to articles by anyone who isn't affiliated with a subscribing university.
What services do modern journals provide?
Historically, when information was spread exclusively via ink on paper, academic journals provided a crucial service to the academic community: distribution.
Over the past few decades, the advent of the internet has fundamentally changed the way we use and share information, but traditional journals remain a staple of academia - not because they continue to contribute unique value, but because they're entrenched in the system and careers still depend on publishing in them.
The internet makes publication of information easy and efficient. Modern academic journals don't provide any additional value that could not be easily and cheaply replicated, and because of their history as a physical medium they have been slow to realize the full potential of the publication of information via the internet. Journals still publish discrete issues, enforce page limits, and sometimes charge to print color figures. In the digital age, all of these practices are unnecessary - laughably so. They also charge exorbitant amounts for access to articles by anyone who isn't affiliated with a subscribing university.
What services do modern journals provide?
- Peer review. Journals don't provide this; volunteer academics do, for free. Journals simply organize it and then profit from it. Any impartial third party could serve the same function and provide their stamp of approval to a paper. There are also numerous alternatives to the current model of peer review employed by journals. One is to use post-publication review and encourage public dialogue on research publication websites. This is already done on journal websites and others like arxiv.
Peer review and some degree of expert filtering is important, but many are frustrated with the current system for good reason. Peer review should have a well-defined and limited scope. Reviewers should check research for soundness and rigor and, where possible, replicate results. This is one of the pillars of the scientific method. Yet, modern peer review typically fails to attempt any sort of replication. It is also subjective and can be driven by political or ideological conflicts instead of validity. Given a medium such as the internet with practically unlimited storage space, reviewers should not be questioning the potential impact of work; they should reject bad science, and nothing else. - Name recognition. You score more points for publishing in some journals than others. Is this a good thing? Papers should be judged by their measurable impact and the ensuing discussion, not which company deemed them worthy of publication. The current system makes it easy to scan someone's list of publications and instantly judge their quality as a researcher. Maybe it shouldn't be so easy.
- Topical organization. I love reading a few specific journals, like Global Ecology and Biogeography - it's a topic that interests me and so I find most of the papers within it interesting. But publishing separate journals for different subjects is hardly necessary. The internet already knows how to categorize and organize information. Wikipedia, Google, and Reddit are three very different examples of how this has been done successfully.
- Formatting and typesetting. This is not a concern with content published online - just use markdown or HTML. Even if a physical printout is required, thanks to tools like LaTeX, these tasks are simpler than ever.
Academic journals provide marginal, replaceable value to the dissemination of research, and by doing so they somehow earn the right to profit from and control access to the results of research (often publicly funded research) they had no part in. This is astonishingly unethical, and more people need to start challenging the status quo like Aaron did.
The rising generation that grew up in the age of the internet believes strongly that information should be free and available, not guarded for profit. Something needs to be done to disrupt academic publishers. The results of scientific research should be freely available to everyone, and private publishers should not act as gatekeepers to knowledge.
I reject the idea that any corporation can profit by publishing publicly funded research that should inherently be free. I want to see widespread rejection of this idea. Share your PDFs. They can't arrest all of us.
Aaron Swartz's Open Access Manifesto
I reject the idea that any corporation can profit by publishing publicly funded research that should inherently be free. I want to see widespread rejection of this idea. Share your PDFs. They can't arrest all of us.
Aaron Swartz's Open Access Manifesto
Saturday, December 29, 2012
What incentives are there to maintain software in academia?
Just read an article in PLoS Comp. Bio. called "Ten Simple Rules for the Open Development of Scientific Software" by Andreas Prlić which was linked by Karthik Ram. A Twitter discussion followed, in which 140 characters was not enough to be sufficiently expressive. Let me start off by saying that I think this was a fantastic article. I'm 100% in agreement and think that these are some important points to make. I start with this caveat because I'm about to dwell on one suggestion that I had a negative reaction to.
From rule 10, "science counts:"
The author hits on the unfortunate practical reality that time spent on software development that doesn't result in widely-recognized deliverables such as publications or grants is essentially time wasted, and will be inversely correlated with your chances of success as an academic.
The troubling part is that this is an extraordinarily short-sighted view of the value of software. Outside of academia, large communities of developers frequently and happily contribute to open source projects for which they receive no tangible benefit. The rewards developers receive vary from education and experience to networking and recognition to simply having fun. Sometimes extrinsic rewards eventually present themselves, and beyond a certain level of growth money becomes increasingly necessary to keep a large project going (see the "Money" chapter from Producing Open Source Software.) Still, popular open source projects such as Linux and Python have value that far outweigh the modest amounts of money that have been funneled into them, and they're still developed largely by unpaid (sometimes anonymous) volunteers.
Scientific software is important, and even very specialized software should be more widely available and used more often. Replication is one of the cornerstones of the scientific method. I envision a future where results and figures from papers are easily replicable upon publication and where people (reviewers especially) are in the habit of checking each others' work. This is already being done on small scales - see Weecology on GitHub for some excellent examples. The problem is this: a scientist who develops code for a single analysis and makes their code publicly available is doing it to benefit the broader scientific community. But code rots over time. Inevitably, when code makes the jump from a single user to many, problems will be discovered. Thus, the benefit provided by open source software is directly related to the effort spent responding to users and maintaining code. And for most projects, this effort has a very low probability of providing the author of the code with an additional grant or publication, so there's little incentive to do it. (There are notable counterexamples - massive projects such as DataONE for which there's already funding for long-term development and maintenance and which tend to result in multiple publications and presentations for those involved.)
So, my question is this: what can be done to provide incentives for the development and maintenance of important scientific code?
From rule 10, "science counts:"
As scientists, the software we write is primarily a means to advance our research and, ultimately, achieve our scientific goals. Whilst the development of software for the consumption of others aligns well with other processes of scientific advancement, it is the science that ultimately counts. Scientific software development fulfils an immediate need, but maintenance of code that is no longer relevant to your own research is a serious time sink, and will rarely lead to your next paper, or secure your next grant or position.
The author hits on the unfortunate practical reality that time spent on software development that doesn't result in widely-recognized deliverables such as publications or grants is essentially time wasted, and will be inversely correlated with your chances of success as an academic.
The troubling part is that this is an extraordinarily short-sighted view of the value of software. Outside of academia, large communities of developers frequently and happily contribute to open source projects for which they receive no tangible benefit. The rewards developers receive vary from education and experience to networking and recognition to simply having fun. Sometimes extrinsic rewards eventually present themselves, and beyond a certain level of growth money becomes increasingly necessary to keep a large project going (see the "Money" chapter from Producing Open Source Software.) Still, popular open source projects such as Linux and Python have value that far outweigh the modest amounts of money that have been funneled into them, and they're still developed largely by unpaid (sometimes anonymous) volunteers.
Scientific software is important, and even very specialized software should be more widely available and used more often. Replication is one of the cornerstones of the scientific method. I envision a future where results and figures from papers are easily replicable upon publication and where people (reviewers especially) are in the habit of checking each others' work. This is already being done on small scales - see Weecology on GitHub for some excellent examples. The problem is this: a scientist who develops code for a single analysis and makes their code publicly available is doing it to benefit the broader scientific community. But code rots over time. Inevitably, when code makes the jump from a single user to many, problems will be discovered. Thus, the benefit provided by open source software is directly related to the effort spent responding to users and maintaining code. And for most projects, this effort has a very low probability of providing the author of the code with an additional grant or publication, so there's little incentive to do it. (There are notable counterexamples - massive projects such as DataONE for which there's already funding for long-term development and maintenance and which tend to result in multiple publications and presentations for those involved.)
So, my question is this: what can be done to provide incentives for the development and maintenance of important scientific code?
Subscribe to:
Posts (Atom)